# Databricks — Evaluating AI Agents Live at the Grounded Reasoning Cup

- Company: Databricks (databricks.com)
- Announced: 2026-08-18T15:00:00+00:00
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://www.databricks.com/blog/evaluating-ai-agents-live-grounded-reasoning-cup
- Record: https://forck.live/items/17499-evaluating-ai-agents-live-at-the-grounded-reasoning-cup
- Subject: Mosaic AI

The Grounded Reasoning Cup challenged 11 academic teams to apply agents developed on OfficeQA Pro to OfficeQA Pro V2, a newly released benchmark built from approximately 120,000 pages of U.S. Treasury documents. Results showed that generalization cannot be assumed. Approaches developed on a familiar benchmark did not always transfer reliably to a new corpus, and out-of-the-box frontier agents averaged less than 30% accuracy. Stanford’s winning team achieved 63.3% accuracy through an end-to-end agent optimization strategy that combined a library of reusable skills, targeted document-representation fallbacks, and adaptive verification. This year, Databricks hosted the inaugural Grounded Reasoning Cup, a first-of-its-kind live AI competition to evaluate AI agents’ ability to reason over complex, enterprise-style document collections. By testing agents on a newly released corpus under live competition conditions, the Grounded Reasoning Cup was designed to help answer one of the hardest questions in AI evaluation: how well do performance improvements on a benchmark generalize to similar, real-world tasks? The competition brought together 11 top academic teams from across the U.S. …

---

Record: https://forck.live/items/17499-evaluating-ai-agents-live-at-the-grounded-reasoning-cup
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
