# Hugging Face — 📚 3LM: A Benchmark for Arabic LLMs in STEM and Code

- Company: Hugging Face (huggingface.co)
- Announced: 2025-08-01T14:25:21+00:00
- Category: research-paper
- Subject: Platform
- Source: https://huggingface.co/blog/tiiuae/3lm-benchmark
- Record: https://forck.live/items/1645-3lm-a-benchmark-for-arabic-llms-in-stem-and-code

Hugging Face announces the release of 3LM, a benchmark for evaluating Arabic large language models on STEM and code generation, consisting of three datasets: Native STEM, Synthetic STEM, and translated code benchmarks (HumanEval-ar, MBPP-ar). The post reports evaluation results for over 40 models.

## Evidence

Verbatim from https://huggingface.co/blog/tiiuae/3lm-benchmark:

> To address this gap, we introduce 3LM (علم), a multi-component benchmark tailored to evaluate Arabic LLMs on STEM (Science, Technology, Engineering, and Mathematics) subjects and code generation.

---

Record: https://forck.live/items/1645-3lm-a-benchmark-for-arabic-llms-in-stem-and-code
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
