# Hugging Face — BenchMIRT: What are LLM benchmarks actually measuring?

- Company: Hugging Face (huggingface.co)
- Announced: 2026-09-01T21:39:07+00:00
- Category: research-paper
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://huggingface.co/blog/allenai/benchmirt
- Record: https://forck.live/items/7290-benchmirt-what-are-llm-benchmarks-actually-measuring
- Subject: Platform

BenchMIRT is a new method for auditing LLM benchmarks at the level of individual prompts, using multidimensional Item Response Theory to separate multiple capabilities (e.g., safety and general reasoning) that contribute to benchmark scores.

## Evidence

Verbatim from https://huggingface.co/blog/allenai/benchmirt:

> Today we’re introducing BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts—the questions and tasks a model is scored on.

---

Record: https://forck.live/items/7290-benchmirt-what-are-llm-benchmarks-actually-measuring
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
