Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Alibaba announces Qwen-VLA, a general-purpose Vision-Language-Action model built on the Qwen multimodal backbone, designed to extend visual perception, language understanding, and spatial reasoning into continuous action generation and trajectory prediction for embodied intelligence tasks such as robotic manipulation and vision-language navigation.
From the source
Qwen-VLA is a general-purpose Vision-Language-Action model. Built upon the Qwen multimodal backbone, it extends visual perception, language understanding, and spatial reasoning into continuous action generation and trajectory prediction.
qwen.ai