# Hugging Face — The Agent Said It Was Done. The Database Disagreed.

- Company: Hugging Face (huggingface.co)
- Announced: 2026-10-03T22:46:38+00:00
- Category: research-paper
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://huggingface.co/blog/microsoft/thinkingbox
- Record: https://forck.live/items/15871-the-agent-said-it-was-done-the-database-disagreed
- Subject: Platform
- Models affected: Claude Opus 5.5, Claude Opus 5, Kimi-K3, GPT-6-Astra, GLM-5.1, Kimi-K2.6, DeepSeek-V4-Pro, Claude Opus 4.6

Microsoft and Hugging Face introduced ThinkingBox, a benchmark that evaluates AI agents on the terminal backend state and side effects they leave behind rather than on tool call correctness. Across 507 stateful business workflows run 20 times each against various LLMs, 79,853 of 121,680 trials failed executable checks, with 67.24% of failures appearing clean despite wrong field values (77.61%), unintended extra effects (43.30%), or missing required effects (25.36%). The benchmark reveals large consistency gaps: Claude Opus 5.5 leads overall at 67.16% pass@1, while Kimi-K3 solves 93.89% of tasks at least once but only 13.41% in all 20 attempts.

## Evidence

Verbatim from https://huggingface.co/blog/microsoft/thinkingbox:

> The gap is substantial. In a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. Executable checks nevertheless found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%.

---

Record: https://forck.live/items/15871-the-agent-said-it-was-done-the-database-disagreed
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
