Apple researchers introduce RLTL;DR, a method where after each failed attempt, the policy sees the verifier outputs and writes its own TL;DR insight, conditioning subsequent rollouts on all previous insights. On challenging tool-calling and coding datasets where standard GRPO training of a Qwen 3.5 9B Thinking policy achieves 0–1% Pass@1, RLTL;DR achieves 14–31% Pass@1 with insights in context during training and 12–13% when no insight is in context at evaluation. The paper also presents SFTL;DR, training only on (task, insight) tuples, which recovers almost full performance from only 4k tuples.
Lead story
Top stories
Models & availability
Latest