Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
OpenAI introduces SimpleQA, a factuality benchmark for evaluating language models on short fact-seeking questions.
From the source
A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.
openai.com