# OpenAI — Why we no longer evaluate SWE-bench Verified

- Company: OpenAI (openai.com)
- Announced: 2026-02-23T11:00:00+00:00
- Subject: GPT / ChatGPT / API
- Source: https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified
- Record: https://forck.live/items/264-why-we-no-longer-evaluate-swe-bench-verified

OpenAI explains that they no longer evaluate SWE-bench Verified due to contamination and flawed tests, and they recommend SWE-bench Pro.

## Evidence

Verbatim from https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified:

> SWE-bench Verified is increasingly contaminated and mismeasures frontier coding progress.

---

Record: https://forck.live/items/264-why-we-no-longer-evaluate-swe-bench-verified
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
