# Hugging Face — IBM and UC Berkeley Diagnose Why Enterprise Agents Fail Using IT-Bench and MAST

- Company: Hugging Face (huggingface.co)
- Announced: 2026-02-18T16:15:45+00:00
- Category: research-paper
- Subject: Platform
- Models affected: Gemini-3-Flash, Kimi-K2, GPT-OSS-120B
- Source: https://huggingface.co/blog/ibm-research/itbenchandmast
- Record: https://forck.live/items/1545-ibm-and-uc-berkeley-diagnose-why-enterprise-agents-fail-using-it-bench-and-mast

IBM Research and UC Berkeley collaborated to study how agentic LLM systems fail in enterprise IT automation tasks. They used MAST (Multi-Agent System Failure Taxonomy) to analyze execution traces from ITBench, a benchmark for SRE, Security, and FinOps automation. They annotated 310 ITBench SRE traces across three models: Gemini-3-Flash, Kimi-K2, and GPT-OSS-120B. Key findings include that frontier models fail cleanly with few failure modes, while large open models suffer cascading failures; incorrect verification is a strong predictor of failure across models; and Kimi-K2 shows spikes in premature termination and unawareness of termination conditions.

## Evidence

Verbatim from https://huggingface.co/blog/ibm-research/itbenchandmast:

> By leveraging MAST to analyze ITBench—the industry benchmark for SRE, Security, and FinOps automation—we turned raw execution traces into structured failure signatures, revealing exactly what broke and how to fix it.

---

Record: https://forck.live/items/1545-ibm-and-uc-berkeley-diagnose-why-enterprise-agents-fail-using-it-bench-and-mast
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
