# Sakana AI — Reinforcement Learning Teachers of Test Time Scaling

- Company: Sakana AI (sakana.ai)
- Announced: 2025-06-22T15:00:00+00:00
- Category: research-paper
- Subject: Sakana models
- Open weights: yes
- Models affected: DeepSeek R1, RLT
- Source: https://sakana.ai/rlt
- Record: https://forck.live/items/4713-reinforcement-learning-teachers-of-test-time-scaling

Sakana AI introduces Reinforcement-Learned Teachers (RLTs), a new method for training teacher models to generate explanations for student models by learning to teach rather than solve problems. The teacher is rewarded based on the student's understanding. Small 7B parameter RLTs outperform much larger models like DeepSeek R1 (671B) in teaching reasoning skills. The code, paper, and open models are released.

## Evidence

Verbatim from https://sakana.ai/rlt:

> Our compact teachers with only 7B parameters are better at teaching reasoning skills than orders-of-magnitude larger LLMs, making advanced AI more affordable and much faster to train.

---

Record: https://forck.live/items/4713-reinforcement-learning-teachers-of-test-time-scaling
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
