From the source
Repetitive degeneration [1] is a common failure mode during inference: the model emits a span (often something like "Wait, let me reconsider…"), then repeats the same span again and again, until the context window is exhausted.
We call this phenomenon the 'doom loop'.
Small reasoning models are more prone to this behavior, especially on long thinking traces and hard problems [2].
The commonly applied inference-time fix is to apply repetition_penalty to reweight the output distribution.
However, this is a band-aid solution and can degrade performance.
Reinforcement learning can target repetitive looping, but it typically requires carefully calibrated rewards and costly online rollouts.
Our method takes a more targeted approach.
We identify the exact token that begins a loop, train the model to prefer coherent alternatives at that single position, and leave the rest of the distribution largely untouched.
The method adapts Antislop [3], training on chosen/rejected pairs that represent a single completion token, using Final Token Preference Optimization (FTPO).
We call our approach “Antidoom”.
…




