
Recursive Language Models — Alex Zhang, MIT PhD
About this episode
From the show’s notesFrom GPU kernels and KernelBench to Recursive Language Models, agent harnesses, and massive multi-agent swarms, Alex Zhang is exploring how much capability we’re leaving on the table by wrapping increasingly powerful models in primitive systems. In this episode, the MIT researcher behind RLMs joins swyx to unpack why Claude Code, Codex, and most coding agents are more similar than they look, why better harness design could unlock capabilities already latent in today’s models, and what happens when you can throw 10,000 agents and tens of millions of dollars at a single problem.
We go deep on GPU Mode and AI-written kernels, research taste and why academics should take bets industry labs won’t, GEV and alternatives to the standard autoregressive language model, and the idea of harnesses as compositional generalizers. Alex explains RLMs, context offloading, programmatic subagent calling, Prime Agent, persistent subagents, and why the “language model” of the future may actually be an invisible swarm of agents underneath a simple interface. We also discuss OpenAI’s massive agent experiments, Kimi swarms, open-ended research at Sakana AI, speculative programmatic tool calling, capability overhang, Neuralese, and where Alex thinks the next big research opportunities may lie.
Read the show’s notes in full
We discuss: • Why AI-generated GPU kernels still leave substantial room for human expertise • How one expert insight can potentially replace enormous amounts of brute-force token search • Why PhD students should take research bets that initially look trivial, weird, or pointless • What SWE-bench, RLMs, ReAct, and Quiet-STaR reveal about research taste • GEV and why a language model does not have to mean an autoregressive text-to-text decoder • Why Claude Code, Codex, Pi, and many modern agent harnesses are structurally very similar • How harness design can improve compositional generalization across tasks and domains • RLMs: context offloading, code execution, recursive subagents, and shared memory • Prime Agent, continual harnesses, and persistent agent-to-agent communication • Why the model you query in the future may secretly be an entire swarm or scaffold • OpenAI’s 10,000-agent experiment, 130B output tokens, and ~$40M-equivalent problem solving • Why much of an agent swarm may be wasted search — and why convergence is still hard • Kimi versus OpenAI’s approach to multi-agent systems • Open-endedness, Sakana AI, and the problem of finding hidden gems in enormous amounts of generated work • Why current frontier models may already have a large capability overhang • Speculative programmatic tool calling and overlapping tool execution with generation • Whether English, code, or an entirely new “Neuralese” constrains how models reason • AI for science, fast-moving benchmarks, and how Alex chooses what research problems to bet on





