From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Images, video and audio represent different aspects of the underlying reality.
Training across these modalities lets us build on a much broader source of data than action demonstrations alone, resulting in more generalization.
We then adapt this foundation through joint video-action training and finetuning for a target embodiment and its corresponding action space.
In robotics, the fine-tuning recipe is just as important as the weights.
We are publishing this report along with the weights to make this process transparent; from the pretraining and midtraining phases to finetuning and inference optimizations for action prediction - where efficiency in particular is a critical and necessary capability for local deployments.
In addition, we analyze hybrid systems that combine fast action prediction with the planning capabilities of frontier reasoning models - and find that fast control makes embodied reasoning more cost- and time efficient - up to 53.64% more success per dollar.
The focus of this report is action prediction applied to robotics - but action prediction extends to digital environments as well.
…