# Apple — RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback

- Company: Apple (apple.com)
- Announced: 2026-10-01
- Category: research-paper
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://machinelearning.apple.com/research/rltl-dr-self-improvement
- Record: https://forck.live/items/15591-rltl-dr-self-improvement-by-internalizing-self-generated-feedback
- Subject: Machine Learning Research
- Models affected: Qwen 3.5 9B

Apple researchers introduce RLTL;DR, a method where after each failed attempt, the policy sees the verifier outputs and writes its own TL;DR insight, conditioning subsequent rollouts on all previous insights. On challenging tool-calling and coding datasets where standard GRPO training of a Qwen 3.5 9B Thinking policy achieves 0–1% Pass@1, RLTL;DR achieves 14–31% Pass@1 with insights in context during training and 12–13% when no insight is in context at evaluation. The paper also presents SFTL;DR, training only on (task, insight) tuples, which recovers almost full performance from only 4k tuples.

## Evidence

Verbatim from https://machinelearning.apple.com/research/rltl-dr-self-improvement:

> On challenging tool-calling and coding datasets (filtered to Pass@128 = 0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14–31% with insights in context during training and, crucially, 12–13% when no insight is in context at eval time.

---

Record: https://forck.live/items/15591-rltl-dr-self-improvement-by-internalizing-self-generated-feedback
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
