From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Experimental draft model speeds up VLM inference on edge and GPU.
Liquid AI released an experimental DSpark draft model for its vision-language model LFM2.5-VL-3B, adding a speculative decoding path that increases memory footprint by 8.9% (280M parameters) while achieving decode speedups up to 3.13x on device and 2.66x on an H100, with end-to-end gains up to 2.62x and 2.27x.
The drafter uses a simplified attention-only architecture with 4 layers and a block size of 9, and ships with day-one support for llama.cpp, MLX-VLM, and SGLang.
The model is open-weight and available in Safetensors and GGUF formats.
From the source
Faster inference: decode speedups up to 3.13x on device and 2.66x on an H100, with end-to-end gains up to 2.62x and 2.27x.
huggingface.co
Reported by