# Hugging Face — Up to 3.2x Faster Inference with LFM2.5-DSpark

- Company: Hugging Face (huggingface.co)
- Announced: 2026-08-20T16:52:57+00:00
- Category: infrastructure-release
- Subject: Platform
- Open weights: yes
- Models affected: LFM2.5-1.2B-Instruct, LFM2.5-2.6B, LFM2.5-8B-A1B, LFM2.5-2.6B-DSpark, LFM2.5-1.2B-Instruct-DSpark, LFM2.5-8B-A1B-DSpark
- Source: https://huggingface.co/blog/LiquidAI/lfm25-dspark
- Record: https://forck.live/items/4245-up-to-3-2x-faster-inference-with-lfm2-5-dspark

Liquid AI releases DSpark draft model checkpoints for three LFM2.5 family models, providing speculative decoding for faster inference (up to 3.18x throughput improvement on GPU, up to 2.87x on-device) and reducing function-calling latency by 57% for LFM2.5-2.6B. The draft models are around 300M parameters each and are available in Safetensors and GGUF formats with day-one support for llama.cpp and SGLang. The high-level claim is "Up to 3.2x Faster Inference" overall, but the specific measured improvements are 3.18x (GPU) and 2.87x (on-device). Quality parity is maintained under greedy decoding, and the emitted sequence is identical to baseline greedy by construction.

## Evidence

Verbatim from https://huggingface.co/blog/LiquidAI/lfm25-dspark:

> Today, we release DSpark draft model checkpoints for three models from our LFM2.5 family: LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B.

---

Record: https://forck.live/items/4245-up-to-3-2x-faster-inference-with-lfm2-5-dspark
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
