Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
OpenAI published an analysis highlighting problems with the SWE-Bench Pro coding benchmark, questioning its reliability and accuracy for evaluating AI models.
From the source
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.
openai.com