Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face announces the release of 3LM, a benchmark for evaluating Arabic large language models on STEM and code generation, consisting of three datasets: Native STEM, Synthetic STEM, and translated code benchmarks (HumanEval-ar, MBPP-ar). The post reports evaluation results for over 40 models.
From the source
To address this gap, we introduce 3LM (ุนูู ), a multi-component benchmark tailored to evaluate Arabic LLMs on STEM (Science, Technology, Engineering, and Mathematics) subjects and code generation.
huggingface.co