Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Alibaba introduces Qwen-RobotWorld, a dual-stream diffusion world model that uses natural language as a universal action interface to unify 20+ robot embodiments and 500+ action categories, enabling cross-scenario physical generalization across manipulation, driving, and navigation. The model employs Qwen2.5-VL as an action encoder within an MMDiT architecture, and is trained on the Embodied World Knowledge (EWK) dataset of 8.6M video-text pairs. It also supports multi-view geometrically consistent generation and human-to-robot transfer via Scene2Robot.
From the source
Qwen-RobotWorld bridges this gap by treating natural language as a universal action interface. ... Language unifies the action space : world knowledge and embodied knowledge reinforce each other within a single model, enabling cross-scenario, cross-task physical generalization .
qwen.ai