# Hugging Face — Inside VAKRA: Reasoning, Tool Use, and Failure Modes of Agents

- Company: Hugging Face (huggingface.co)
- Announced: 2026-04-15T12:07:25+00:00
- Category: research-paper
- Subject: Platform
- Source: https://huggingface.co/blog/ibm-research/vakra-benchmark-analysis
- Record: https://forck.live/items/1519-inside-vakra-reasoning-tool-use-and-failure-modes-of-agents

Hugging Face and IBM Research published a blog post introducing VAKRA, a tool-grounded, executable benchmark for evaluating AI agents' reasoning and tool use in enterprise environments. The post describes the benchmark's four tasks, dataset details, and analysis of failure modes observed on different tasks.

## Evidence

Verbatim from https://huggingface.co/blog/ibm-research/vakra-benchmark-analysis:

> We recently introduced VAKRA, a tool-grounded, executable benchmark for evaluating how well AI agents reason and act in enterprise-like environments.

---

Record: https://forck.live/items/1519-inside-vakra-reasoning-tool-use-and-failure-modes-of-agents
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
