Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
This blog introduces HELMET, a comprehensive benchmark for evaluating long-context language models (LCLMs) that addresses limitations of existing evaluations by providing diverse, controllable, and reliable metrics. The benchmark is presented at ICLR 2025.
From the source
In this work, we propose HELMET (How to Evaluate Long-Context Models Effectively and Thoroughly), a comprehensive benchmark for evaluating LCLMs that improves upon existing benchmarks in several ways— diversity, controllability, and reliability.
huggingface.co