Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Alibaba announces Qwen-RobotManip, a generalizable Vision-Language-Action (VLA) foundation model built upon Qwen-VL, which introduces a unified alignment framework across representation, motion, and behavioral dimensions for robotic manipulation. The model uses only open-source robotic manipulation datasets and human demonstration videos to construct a ~38,100 hours pretraining corpus, and demonstrates emergent generalization capabilities across various real-robot platforms and tasks.
From the source
Qwen-RobotManip is a generalizable Vision-Language-Action (VLA) foundation model built upon Qwen-VL. It introduces a unified alignment framework across the representation, motion, and behavioral dimensions of manipulation, making large-scale multi-source training coherent rather than conflicting. Using only open-source robotic manipulation datasets and human demonstration videos without any proprietary data collection, Qwen-RobotManip constructs a ~38,100 hours pretraining corpus and already exhibits emergent generalization capabilities.
qwen.ai