Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
IBM Research and UC Berkeley collaborated to study how agentic LLM systems fail in enterprise IT automation tasks. They used MAST (Multi-Agent System Failure Taxonomy) to analyze execution traces from ITBench, a benchmark for SRE, Security, and FinOps automation. They annotated 310 ITBench SRE traces across three models: Gemini-3-Flash, Kimi-K2, and GPT-OSS-120B. Key findings include that frontier models fail cleanly with few failure modes, while large open models suffer cascading failures; incorrect verification is a strong predictor of failure across models; and Kimi-K2 shows spikes in premature termination and unawareness of termination conditions.
From the source
By leveraging MAST to analyze ITBench—the industry benchmark for SRE, Security, and FinOps automation—we turned raw execution traces into structured failure signatures, revealing exactly what broke and how to fix it.
huggingface.co