Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Liquid AI releases DSpark draft model checkpoints for three LFM2.5 family models, providing speculative decoding for faster inference (up to 3.18x throughput improvement on GPU, up to 2.87x on-device) and reducing function-calling latency by 57% for LFM2.5-2.6B. The draft models are around 300M parameters each and are available in Safetensors and GGUF formats with day-one support for llama.cpp and SGLang. The high-level claim is "Up to 3.2x Faster Inference" overall, but the specific measured improvements are 3.18x (GPU) and 2.87x (on-device). Quality parity is maintained under greedy decoding, and the emitted sequence is identical to baseline greedy by construction.
From the source
Today, we release DSpark draft model checkpoints for three models from our LFM2.5 family: LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B.
huggingface.co