Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Tencent Hunyuan introduces CL-bench Life, a benchmark for evaluating context learning ability in real-life settings, containing 405 context-task pairs and 5,348 human-written rubrics. Evaluations of 12 language models show they solve only 14.5% of tasks on average, with the best model (GPT-5.5 High) solving 22.2%.
From the source
We introduce CL-bench Life, a rigorous benchmark for evaluating context learning ability in real-life settings and guiding future model development.
hunyuan.tencent.com