From the source
Liquid AI released DSpark draft model checkpoints for three LFM2.5 models, enabling speculative decoding that achieves up to 3.18x throughput improvement on GPU and up to 2.87x on-device without changing output quality.
The draft models are available on Hugging Face, with integrations open-sourced in llama.cpp and SGLang.




