Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Tencent Hunyuan introduces CL-bench, a benchmark to evaluate whether language models can learn new knowledge from context rather than relying solely on parametric memory, covering four types of real-world context learning scenarios with a contamination-free design.
From the source
To evaluate how far current models are from being true context learners, we built CL-bench, a benchmark that tests whether LMs can learn new knowledge from context and apply it correctly.
hunyuan.tencent.com