From the source
The blog post investigates why MMLU evaluation scores on the Open LLM Leaderboard differ from those reported in the LLaMA paper, finding that different implementations of the same benchmark yield varying results and even change model rankings.


