Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
OpenAI introduces PaperBench, a benchmark to evaluate AI agents' ability to replicate state-of-the-art AI research.
From the source
We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research.
openai.com