# Hugging Face — Accelerating vision-language models with LFM2.5-VL-DSpark

- Company: Hugging Face (huggingface.co)
- Announced: 2026-09-24T14:08:57+00:00
- Category: model-update
- Coverage: 1 outlet
- Announcement: yes
- Group: models
- Source: https://huggingface.co/blog/LiquidAI/lfm2-5-vl-dspark
- Record: https://forck.live/items/13710-accelerating-vision-language-models-with-lfm2-5-vl-dspark
- Subject: Platform
- Open weights: yes
- Models affected: LFM2.5-VL-3B, LFM2.5-VL-3B-DSpark

Liquid AI released an experimental DSpark draft model for its vision-language model LFM2.5-VL-3B, adding a speculative decoding path that increases memory footprint by 8.9% (280M parameters) while achieving decode speedups up to 3.13x on device and 2.66x on an H100, with end-to-end gains up to 2.62x and 2.27x. The drafter uses a simplified attention-only architecture with 4 layers and a block size of 9, and ships with day-one support for llama.cpp, MLX-VLM, and SGLang. The model is open-weight and available in Safetensors and GGUF formats.

## Evidence

Verbatim from https://huggingface.co/blog/LiquidAI/lfm2-5-vl-dspark:

> Faster inference: decode speedups up to 3.13x on device and 2.66x on an H100, with end-to-end gains up to 2.62x and 2.27x.

## Around this story

Outlets this record can name and link:

- Marktechpost — Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding: https://marktechpost.com/2026/09/25/liquid-ai-releases-lfm2-5-vl-3b-dspark-speculative-decoding-for-vision-language-models-with-up-to-3-13x-faster-decoding

---

Record: https://forck.live/items/13710-accelerating-vision-language-models-with-lfm2-5-vl-dspark
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
