From the source
Cohere's research shows that multilingual models can be trained to reason in the user's language, not just answer in it, through a data mixing strategy built on three pillars: broader language coverage, a small share of multilingual non-reasoning data, and English reasoning data as a backbone.
The resulting model, Tiny Aya L2-Thinker (3.35B), reasons in the prompt language over 93% of the time across 60 languages with minimal accuracy loss, and the weights and data are released.
The research also finds that inference-time language forcing approaches are ineffective, so in-language reasoning must come from training.
