# Hugging Face — Fixing Open LLM Leaderboard with Math-Verify

- Company: Hugging Face (huggingface.co)
- Announced: 2025-02-14T00:00:00+00:00
- Category: capability-change
- Subject: Platform
- Open weights: yes
- Models affected: Qwen, DeepSeek
- Source: https://huggingface.co/blog/math_verify_leaderboard
- Record: https://forck.live/items/1747-fixing-open-llm-leaderboard-with-math-verify

Hugging Face re-evaluated all 3,751 models on the Open LLM Leaderboard using the Math-Verify tool to fix math evaluation issues, resulting in an average 4.66-point score increase and significant reshuffling of rankings, especially for Qwen and DeepSeek models.

## Evidence

Verbatim from https://huggingface.co/blog/math_verify_leaderboard:

> we've used Math-Verify to thoroughly re-evaluate all 3,751 models ever submitted to the Open LLM Leaderboard, for even fairer and more robust model comparisons!

---

Record: https://forck.live/items/1747-fixing-open-llm-leaderboard-with-math-verify
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
