From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
RLTL;DR method enables self-improvement from failed attempts via self-generated feedback.
Apple researchers introduce RLTL;DR, a method where after each failed attempt, the policy sees the verifier outputs and writes its own TL;DR insight, conditioning subsequent rollouts on all previous insights.
On challenging tool-calling and coding datasets where standard GRPO training of a Qwen 3.5 9B Thinking policy achieves 0–1% Pass@1, RLTL;DR achieves 14–31% Pass@1 with insights in context during training and 12–13% when no insight is in context at evaluation.
The paper also presents SFTL;DR, training only on (task, insight) tuples, which recovers almost full performance from only 4k tuples.
From the source
On challenging tool-calling and coding datasets (filtered to Pass@128 = 0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14–31% with insights in context during training and, crucially, 12–13% when no insight is in context at eval time.
machinelearning.apple.com