From the source
It upgrades a pre-trained model's tokenizer in place , without retraining from scratch.
We doubled the vocabulary from 65K to 128K to fix the languages our original tokenizer split too finely.
Hindi and Vietnamese now take roughly 2.4× and 2.6× fewer tokens, and Thai up to 4.0× fewer, which we estimate makes per-character decoding 2.2 to 3.7× faster for these languages on-device.
The quality of the languages the model already handled well holds steady.
LFM2.5-8B-A1B is available on Hugging Face , and the expanded 128K tokenizer is released alongside it.
The full method, ablations, and per-device numbers are in the technical report .
Why on-device tokenizers stay small A tokenizer is fixed at the start of pre-training, and it splits the vocabulary in proportion to whatever the training corpus looked like then.
Languages that were underrepresented at that point get broken into far more tokens per word than the ones the tokenizer was tuned for [1, 2].
Because a model runs its decoder once per output token, that extra fragmentation costs real latency, compute, and energy for users of those languages.
…




