Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face introduces a benchmark and harness for evaluating how efficiently open models use the transformers library in agentic tasks, measuring not just correctness but also cost, latency, and token usage across different library configurations.
From the source
We wanted the whole process instead: not just whether the agent got it right, but how much work it took to get there, and how that shifts across models, library revisions, and tasks.
huggingface.co