# Thinking Machines Lab — On-Policy Distillation

- Company: Thinking Machines Lab (thinkingmachines.ai)
- Announced: 2025-10-27
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://thinkingmachines.ai/blog/on-policy-distillation/
- Record: https://forck.live/items/18568-on-policy-distillation
- Subject: Thinking Machines / Inkling

LLMs are capable of expert performance in focused domains, a result of several capabilities stacked together: perception of input, knowledge retrieval, plan selection, and reliable execution. This requires a stack of training approaches, which we can divide into three broad stages Pre-training teaches general capacities such as language use, broad reasoning, and world knowledge. Mid-training imparts domain knowledge, such as code, medical databases, or internal company documents. Post-training elicits targeted behavior, such as instruction following, reasoning through math problems, or chat. Smaller models with stronger training often outperform larger, generalist models in their trained domains of expertise. There are many benefits to using smaller models: they can be deployed locally for privacy or security considerations, can continuously train and get updated more easily, and save on inference costs. Taking advantage of these requires picking the right approach for the later stages of training. Approaches to post-training a “student” model can be divided into two kinds: On-policy training samples rollouts from the student model itself, and assigns them some reward. …

---

Record: https://forck.live/items/18568-on-policy-distillation
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
