From the source
AI to simulate both the visuals and sounds of the world, all in real-time Oliver Cameron May 17th, 2026 World models are a new form of generative intelligence.
Unlike language models, which learn from text, world models learn directly from the world itself through the pixels, motion, and actions encoded in large-scale video.
In the process, these models become capable of understanding and simulating an approximation of the world in real-time.
Today, we’re excited to share a preview of Starchild-1, the world’s first multimodal world model.
In Starchild-1, we’ve trained a general world model that autoregressively generates synchronized audio and video in real-time, while continuously responding to streaming user input.
Starchild-1 goes beyond traditional world models, which have been limited to learning and generating visuals alone, with no sound.
Audio-video jointly generated in real-time Beyond Visual Simulation The world is not silent.
It's full of conversation, emotion, crashing waves, and chirping birds.
Sound is a rich signal about how the world works, and humans use it constantly to understand and explore the world around them.
Machines should too.
…





