
Runway’s Bet Beyond Video: World Models, Robotics, and the Neural OS — Anastasis Germanidis
About this episode
From the show’s notesFrom making one of the earliest bets on generative video to building world models that simulate physics, power robots, and render software interfaces directly as pixels, Runway is pushing far beyond AI video generation. In this episode, Runway co-founder Anastasis Germanidis joins swyx and Vibhu to unpack how the company went from experimental creative tools to frontier video models — and why Runway now sees video generation as a path toward general-purpose world simulation.
We go deep on the evolution from Gen-1 and Gen-2 to Gen-3 and Gen-4.5, including the thousand-A100 bet that helped kickstart Runway’s video research, the surprising weekend hack behind Gen-2, and the existential moment when OpenAI released Sora and people declared “Runway’s done.” Anastasis explains why scaling video models may be enough to learn physics, why real-time generation is inevitable, and how Runway is turning those models into simulators for robotics and interactive worlds.
Read the show’s notes in full
We also explore Interface World Models — software rendered directly by neural networks with no HTML, CSS, or React — and the possibility of a fully neural operating system. Anastasis lays out his “Lucid Dream Test” for world models, explains how third-person video could unlock robotics at scale, and discusses video agents, omni models, AI-native creative workflows, open-source world models, and what happens when models learn from more than just text, images, and video.
We discuss: • How Runway went from creative tools to building frontier generative video models • Why Runway made a massive early bet on a cluster of 1,000 A100s • The Stable Diffusion story and Runway’s role in its development • Why text-to-video alone was never enough for professional creative work • How Gen-2 emerged from a weekend experiment combining text, depth, and video • Why camera control helped lead Runway toward world models • Why learning directly from the physical world may unlock capabilities language cannot capture • Whether simply scaling video prediction can teach models physics • What happened inside Runway when OpenAI released Sora • How the team scaled Gen-3 by roughly 10× in a three-month push • Why Anastasis believes real-time video generation is inevitable • Interface World Models and generating software pixels directly instead of generating code • Why the endgame could be a fully neural operating system • How Runway is using world models as simulators for robotics • Why third-person internet video could be the largest training source for robotic intelligence • World Action Models and turning video models into robot policies • The “Lucid Dream Test” for knowing when world models are good enough • Video agents, omni models, and why orchestration eventually moves inside the model • Open-source world models, Chinese video models, and the global race in generative video • Why artists are responding differently as AI tools become more controllable • The future of real-time creative tools, physical AI, multimodal models, and scientific simulation





