From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Training method reduces tool-call failures by 21.2%
How users interact with our products in the real world is a valuable source for model training.
The data is abundant, reflects the actual distribution of user tasks, and captures user corrections and tool failures that synthetic environments may miss.
A common way to learn from real-world data is rejection sampling fine-tuning: judge each session’s outcome, keep the successful ones, and train the model to imitate them.
But a successful outcome does not mean every step was correct, so imitating the whole trajectory risks reinforcing bad intermediate behaviors in addition to good ones.
Discarding unsuccessful sessions also loses critical evidence of where the model falls short.
We combine rejection sampling fine-tuning with hint-guided self-distillation to learn from both successful and unsuccessful sessions.
A hint is a short corrective instruction grounded in information the model already had when it made the mistake.
Useful steps from successful sessions remain imitation targets, while grounded hints turn avoidable mistakes into correction targets.
…
From the source
In live use, the later trained checkpoint reduced tool-call failures by 21.2% relative to an earlier trained checkpoint.
perplexity.ai