Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face introduces the LiveCodeBench leaderboard, a new benchmark for evaluating code LLMs that collects coding problems over time from various coding contest platforms to prevent contamination, and includes four scenarios: code generation, self-repair, code execution, and test output prediction.
From the source
We are excited to introduce the LiveCodeBench leaderboard, based on LiveCodeBench, a new benchmark developed by researchers from UC Berkeley, MIT, and Cornell for measuring LLMs’ code generation capabilities.
huggingface.co