Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face re-evaluated all 3,751 models on the Open LLM Leaderboard using the Math-Verify tool to fix math evaluation issues, resulting in an average 4.66-point score increase and significant reshuffling of rankings, especially for Qwen and DeepSeek models.
From the source
we've used Math-Verify to thoroughly re-evaluate all 3,751 models ever submitted to the Open LLM Leaderboard, for even fairer and more robust model comparisons!
huggingface.co