Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Sakana AI introduces Reinforcement-Learned Teachers (RLTs), a new method for training teacher models to generate explanations for student models by learning to teach rather than solve problems. The teacher is rewarded based on the student's understanding. Small 7B parameter RLTs outperform much larger models like DeepSeek R1 (671B) in teaching reasoning skills. The code, paper, and open models are released.
From the source
Our compact teachers with only 7B parameters are better at teaching reasoning skills than orders-of-magnitude larger LLMs, making advanced AI more affordable and much faster to train.
sakana.ai