From the source
It delivers competitive vision performance against models twice its size, while running faster across a range of CPU and GPU deployments even compared with models that have fewer parameters.
LFM2.5-VL-3B builds on our previous LFM2-VL-3B , with significant improvements in screen understanding, grounding, function calling, and multi-image input.
LFM2.5-VL-3B is a non-reasoning model that answers directly, keeping latency low for real-time and on-device applications.
The model is available today on Hugging Face and our Playground .
Check out our docs on how to run and fine-tune it locally.
What’s New LFM2.5-VL-3B extends the vision-language capabilities of our previous release with four major improvements: Screen/UI understanding.
LFM2.5-VL-3B has a strong understanding of digital screens across mobile, web, and desktop.
It averages 80.7 on ScreenSpot-v2, far ahead of the much larger Gemma-4-E4B (51.2) and Qwen 3.5 4B (78.5) and close behind the larger InternVL-3.5-4B (84.1) Function calling.
…




