From the source
It’s a fast, lightweight foundation for developers to fine-tune and deploy in agentic workflows.
Built on the LFM2 architecture, it delivers exceptionally fast inference and runs everywhere, from cloud GPUs to low-cost CPUs (213 tok/s decode speed on Galaxy S25 Ultra, 42 tok/s on a Raspberry Pi 5).
Despite its small size, it’s surprisingly capable at tool use and data extraction tasks.
The base (LFM2.5-230M-Base) and post-trained (LFM2.5-230M) models are available today on Hugging Face .
Check out our docs on how to run and fine-tune them locally.
Training & Fine-tuning The model was pre-trained for 19T tokens, including a 32K context extension phase.
We apply a lightweight post-training recipe designed to preserve flexibility for developers targeting their own downstream applications.
The recipe consists of three stages: (1) supervised fine-tuning with distillation from LFM2.5-350M, (2) direct preference optimization, and (3) multi-domain reinforcement learning .
…




