Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face and IBM Research published a blog post introducing VAKRA, a tool-grounded, executable benchmark for evaluating AI agents' reasoning and tool use in enterprise environments. The post describes the benchmark's four tasks, dataset details, and analysis of failure modes observed on different tasks.
From the source
We recently introduced VAKRA, a tool-grounded, executable benchmark for evaluating how well AI agents reason and act in enterprise-like environments.
huggingface.co