Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Ask your AI
Top stories
Models & availability
Latest
BenchMIRT is a new method for auditing LLM benchmarks at the level of individual prompts, using multidimensional Item Response Theory to separate multiple capabilities (e.g., safety and general reasoning) that contribute to benchmark scores.
From the source
Today we’re introducing BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts—the questions and tasks a model is scored on.
huggingface.co