From the source
Today, we release LFM2.5-VL-450M, an improved version of LFM2-VL-450M with grounding capabilities, better instruction following, and function calling support.
The result is a compact model that turns image streams into structured, actionable outputs in real time, even on edge hardware.
LFM2.5-VL-450M is available on Hugging Face , LEAP , and our Playground .
Check out our docs on how to run and fine-tune it locally.
What’s new Compared to our LFM2-VL-450M that we released a few months ago, we scaled the pre-training of LFM2.5-VL-450M from 10T to 28T tokens, followed by post-training focused on improving multimodal behavior in production settings.
In particular, we used preference optimization and reinforcement learning to improve grounding, instruction following, and overall reliability across vision-language tasks.
Bounding box prediction: 0 → 81.28 on RefCOCO-M We added object detection, allowing the model to identify objects in an image and locate them with bounding boxes.
Improved multilingual image understanding: MMMB 54.29 → 68.09, covering Arabic, Chinese, French, German, Japanese, Korean, Portuguese, Spanish …




