From the source
Today, we release updated 4-bit checkpoints for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B trained with Quantization-Aware Distillation (QAD) .
QAD is a technique to distill a high-precision teacher model into a quantized student model.
These checkpoints keep the low memory footprint and high throughput of Q4_0 GGUFs while recovering most of the accuracy lost to quantization: all four land at roughly 97% of their BF16 averages.
The QAD GGUFs are available today on Hugging Face: LFM2.5-230M , LFM2.5-350M , LFM2.5-1.2B-Instruct , and LFM2.5-2.6B .
Benchmarks For all four models, we compare their released GGUFs produced with post-training quantization (PTQ) against the trained QAD Q4_0 checkpoints on a benchmark suite spanning reasoning, instruction-following, tool use, and agentic capabilities: GPQA Diamond, MMLU-Pro, IFEval, IFBench, Multi-IF, and BFCLv4.
The BF16 GGUF serves as the in-format ceiling.
We also add one scale-appropriate math evaluation: GSM8K for LFM2.5-230M and LFM2.5-350M, and AIME25 for LFM2.5-1.2B-Instruct and LFM2.5-2.6B.
We report the mean across five repeats.
…



