NPHardEval Leaderboard: Unveiling the Reasoning Abilities of Large Language Models through Complexity Classes and Dynamic Updates
Hugging Face introduces the NPHardEval leaderboard, a dynamic benchmark for evaluating LLM reasoning abilities using complexity classes, with 900 algorithmic questions updated monthly.
From the source
We're happy to introduce the NPHardEval leaderboard, using NPHardEval, a cutting-edge benchmark developed by researchers from the University of Michigan and Rutgers University.