From the source
Today, we are releasing OfficeQA Pro V2 , a new benchmark designed to evaluate whether AI agents can generalize to unfamiliar, enterprise-style grounded-reasoning tasks.
Seven months ago, we introduced the OfficeQA benchmark to measure how well AI systems answer analytical questions using evidence from large document collections, an extremely common and important enterprise task that we found agents struggled with.
Since its introduction, OfficeQA, and its frontier subset, OfficeQA Pro , has become an important measure for frontier model and agent capabilities, driving progress in document retrieval, parsing, and analytical reasoning.
But this progress raises a fundamental question: do these improvements reflect broader advances in grounded reasoning, or progress specific to one corpus and task distribution?
This distinction matters in enterprise settings, where agents rarely operate on a single, stable document collection.
Our new benchmark, OfficeQA Pro V2, is designed to test that generalization directly.
…






