Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Distribution-aware speculative decoding (DAS) is a framework that reduces rollout time in RL post-training by up to 50% without altering model outputs. It uses a training-free adaptive suffix tree drafter and a length-aware scheduling strategy to handle stragglers and exploit historical trajectory data.
From the source
Distribution-aware speculative decoding (DAS) is a novel framework that significantly alleviates the rollout bottleneck in RL post-training — delivering up to 50% speedup without touching model outputs.
together.ai