Benchmarks 201: Why Leaderboards > Arenas >> LLM-as-Judge
About this episode
From the show’s notesThe first AI Engineer World’s Fair talks from OpenAI and Cognition are up! In our Benchmarks 101 episode back in April 2023 we covered the history of AI benchmarks, their shortcomings, and our hopes for better ones. Fast forward 1.5 years, the pace of model development has far exceeded the speed at which benchmarks are updated. Frontier labs are still using MMLU and HumanEval for model marketing, even though most models are reaching their natural plateau at a ~90% success rate (any higher and they’re probably just memorizing/overfitting). From Benchmarks to Leaderboards Outside of being stale, lab-reported benchmarks also suffer from non-reproducibility. The models served through the API also change over time, so at different points in time it might return different scores. Today’s guest, Clémentine Fourrier, is the lead maintainer of HuggingFace’s OpenLLM Leaderboard. Their goal is standardizing how models are evaluated by curating a set of high quality benchmarks, and then publishing the results in a reproducible way with tools like EleutherAI’s Harness. The leaderboard was first launched summer 2023 and quickly became the de facto standard for open source LLM performance. To give you a sense for the scale: * Over 2 million unique visitors * 300,000 active community members * Over 7,500 models evaluated Last week they announced the second version of the leaderboard. Why? Because models were getting too good! The new version of the leaderboard is based on 6 benchmarks: * 📚 MMLU-Pro (Massive Multitask Language Understanding - Pro version, paper) * 📚 GPQA (Google-Proof Q&A Benchmark, paper) * 💭MuSR (Multistep Soft Reasoning, paper) * 🧮 MATH (Mathematics Aptitude Test of Heuristics, Level 5 subset, paper) * 🤝 IFEval (Instruction Following Evaluation, paper) * 🧮 🤝 BBH (Big Bench Hard, paper) You can read the reasoning behind each of them on their announcement blog post. These updates had some clear winners and losers, with models jumping up or down up to 50 spots at once; the most likely reason for this is that the models were overfit to the benchmarks, or had some contamination in their training dataset. But the most important change is in the absolute scores. All models score much lower on v2 than they do on v1, which now creates a lot more room for models to show improved performance. On Arenas Another high-signal platform for AI Engineers is the LMSys Arena, which asks users to rank the output of two different models on the same prompt, and then give them an ELO score based on the outcomes. Clémentine called arenas “sociological experiments”: it tells you a lot about the users preference, but not always much about the model capabilities. She pointed to Anthropic’s sycophancy paper as early research in this space: We find that when a response matches a user’s views, it is more likely to be preferred. Moreover, both humans and preference models (PMs) prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time. The other issue is that Arena rankings aren’t reproducible, as you don’t know who ranked what and what exactly the outcome was at the time of ranking. They are still quite helpful as tools, but they aren’t a rigorous way to rank capabilities of the models. Her advice for both arena and leaderboard is to use these tools as ranges; find 3-4 models that fit your needs (speed, cost, capabilities, etc) and then do vibe checks to figure out which one is best for your specific task. LLMs aren’t good judges In the last ~6 months, there has been an increased interest in using LLMs as Judges: rather than asking a person to evaluate the outcome of a model, you can ask a more powerful LLM to score it. We covered this a bit in our Brightwave episode last month as well. HuggingFace also has a cookbook on it, but Clémentine was actually not a fan of this approach: * Mode collapse: if you are asking a model to choose which output is better, it will just self-reinforce its own preferences. It will also prefer models from its own family (i.e. GPT models will prefer other GPT models over Claude outputs). If these outputs are then used to fine-tune the model, you will further mode collapse the model. Cohere for example has said they do not train on any model-generated data to avoid this. * Positional bias: LLMs usually prefer the first answer, so you can’t naively give them options and ask them to rank them, but you also have to mix up the order in which they appear. * Don’t score, rank: rather than asking a model to assign a score to each output, you should have it stack-rank them. The models aren’t trained to score things, so even though they might understand what response is better, assigning a score to it is hard. If you do have to use LLMs as Judges (we aren’t all ScaleAI-rich!), she suggested using an open LLM like Prometheus or JudgeLM to make sure you can reproduce those rankings in the future. Show Notes * Clémentine Fourrier * Hugging Face * OpenLLM v2 Leaderboard * Let’s talk about LLM Evaluation * Leaderboard V2 Blog Post * Latent Space Benchmarks 101 * Gradient AI epsiode on Long Context Evals * Allen AI long context novel evals Companies and Organizations * Anthropic * Cohere * EleutherAI * INRIA * ICLR (International Conference on Learning Representations) People * Aidan Gomez * Dan Hendrycks * Edward Beeching * Haley Sholkoff * Lewis Tunstall * Nathan Habib * Thomas Scialom Projects, Models, and Benchmarks * LMSys Arena * ARC AGI Challenge * Allen Institute ARC Challenge * BigBench * GAIA benchmark * GPQA * GSM 8K * IFEval * LightEval * ML perf * MMLU * JudgeLM * Prometheus * RavenWolf * SWE-Bench * Vantage Timestamps * [00:00:00] Introductions * [00:02:32] How Clémentine went from geology to AI * [00:05:52] Origin of the OpenLLM Leaderboard * [00:09:06] How v1 Benchmarks Were Selected * [00:10:49] The Problem with Current Benchmarks * [00:13:45] Saturating benchmarks and the future of evaluation * [00:16:14] Issues with human evaluations * [00:24:07] AI girlfriends as the multi-turn benchmark * [00:25:35] What's New in OpenLLM leaderboard V2 * [00:28:12] Benchmark Answers Black Market * [00:30:21] The impact of prompt formatting on model evaluation scores * [00:33:30] Difficulty and Computational Constraints of Evals * [00:36:28] The Responsibility of Setting Standards * [00:40:35] The Economics of OpenLLM * [00:44:15] Long context reasoning benchmarks * [00:46:34] Agent benchmarks, GAIA, and the ARC AGI challenge * [00:50:43] Vibe check for benchmarks * [00:53:16] Request for benchmarks * [00:56:48] v3 predictions? Transcript Alessio [00:00:00]: Hey everyone, welcome to the Latent Space podcast. This is Alessio, partner and CTO-in-Residence at Decibel Partners, and I'm joined by my co-host Swyx, founder of Smol AI. Swyx [00:00:13]: Hey, and today we have a super special guest that we've been trying to book on the schedule for a while. It's Clémentine Fourrier. I'm trying my best to do the French, but maybe you can do a better job of it than me. Clémentine [00:00:26]: This was perfect. It's Clémentine Fourrier, but your pronunciation was really on point. Swyx [00:00:31]: There was a Fourrier, which is very sort of French intonation, which I don't really understand. So I'll introduce you off of your LinkedIn and I would love for you to fill in the blanks. You are currently a research scientist at Hugging Face and the maintainer of the OpenLLM leaderboard, which we'll talk about very shortly. Previously, you were at INRIA as well, but then it looks like you also concurrently got your PhD at the same time. How does that work? Is that a very common thing? Clémentine [00:01:01]: So I basically did my PhD at INRIA, technically. So INRIA funded my PhD and PhDs in France are three years, but I also worked as an engineer at INRIA before my PhD, hence maybe the confusion. Swyx [00:01:14]: I think there's a rise in universities having sort of industrial attachments to these things. And I think it actually makes for a much more grounded study, especially if you're doing your sort of graduate studies and all these things. I think it's rising in North America as well with Berkeley and with Waterloo in Toronto. Cool. Like, you know, there's, there's a lot of other things we can, we can introduce. I can't really pronounce the name of the, the university you went to, but what else should people know? Clémentine [00:01:44]: So I actually, technically I'm an engineer in geology. So I studied rocks and I graduated in 2015 after having done like extensive studies about rocks. And I discovered I was very bad at it, but I was very good at computer science. So I went to computer science. What stuck with me though is that geology is very much an experimental science. And I think that machine learning is very much an experimental science too, even though people want to claim that it's pure math. And I worked on several machine learning projects throughout the years, a bit of the prediction of illnesses in the brain at Brain and Spine Institute in Paris. I worked as an engineer in a research team in NLP where I did my thesis and then I joined Hugging Face. Swyx [00:02:32]: Do you have a favorite rock fact or sort of rock story before we get into the NLP stuff? Clémentine [00:02:38]: Okay. I was not expecting this question. Swyx [00:02:43]: I did my geography A-levels and I always loved learning about like isostasy and stuff like that where you have different plates kind of up and down in the mantle. And I don't think people think about vertical dimensions to geographical plates, but it's real. Clémentine [00:03:02]: Yeah, definitely. And like when you do geology, the time scale is just not the same. There is like one specific place in France where you can see rocks that are 1 billion years old and like the sheer scale of this is huge. Yeah, that's what I loved about geology, that the scale is completel





