From the source
Both are bidirectional encoders built on the LFM2 hybrid architecture.
They are designed to be fine-tuned for classification, natural language understanding, and token-level tasks.
On these, they match the quality of larger encoders while scaling much more gently with input length up to a context length of 8,192 tokens.
This keeps document-scale workloads fast on the hardware you already have, including CPU-only environments.
The LFM2.5-Encoder-230M and LFM2.5-Encoder-350M models are available today on Hugging Face.
Check out our docs on how to run and fine-tune them locally.
Why a general-purpose encoder?
Last month, we released LFM2.5-Retrievers , a pair of models built for multilingual retrieval tasks.
LFM2.5-Encoders are close relatives from the same LFM family, but they are built for a broader purpose.
The earlier LFM2.5-Retrievers were trained for retrieval tasks, while the LFM2.5-Encoders are pre-trained with a masked-language objective, so they can be adapted to a range of downstream tasks, including classification, token-level tasks, and retrieval.
…




