
The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
About this episode
From the show’s notesFrom adding GPUs a year before ChatGPT to building the cloud primitives behind elastic inference, agent sandboxes, post-training, and production AI workloads, Modal has quietly become one of the most important infrastructure companies in AI. In this episode, Modal CTO Akshat Bubna joins swyx and Vibhu after Modal’s Series C to unpack why AI applications don’t fit traditional cloud assumptions, why Kubernetes was never designed for bursty compute-heavy workloads, and why Modal is now shifting from developer experience to agent experience.
We go deep on Modal’s AI infra stack: serverless functions, decorator-based infrastructure, elastic inference for custom models, GPU snapshotting, DeFlash, speculative decoding, Auto Endpoints, sandboxes, persistent storage, networked containers, private IPv6, RDMA, multi-node training, and Modal’s capacity pool across 17 cloud providers. Akshat also explains why RL rollouts can require 100,000 sandboxes, why production agents need hard guardrails, why observability may matter more than reading code, and why AI has made infrastructure exciting again.
Read the show’s notes in full
We discuss: • Why Kubernetes wasn’t built for bursty AI workloads • How Modal started as a better runtime before becoming an AI cloud • Why Modal added GPUs a year before ChatGPT • The shift from developer experience to agent experience • Why observability matters when agents are writing the code • Elastic inference for custom models across audio, video, robotics, and comp bio • GPU snapshotting, cold starts, and why inference workloads are so bursty • Why RL rollouts can require 100,000 sandboxes • DeFlash, speculative decoding, and frontier-level inference performance • Auto Endpoints and making optimized inference easier to deploy • What Modal adds beyond vLLM, SGLang, and raw GPU rental • Modal’s 17-cloud capacity pool and “supercloud” strategy • Networked sandboxes, sidecars, private IPv6, and RDMA • Serverless multi-node training for post-training and research workloads • Auto-research, model-guided sweeps, and agents launching GPU experiments • Compute strategy, capacity planning, and batch tiers • Why production agents need specialized sandboxes and hard guardrails • Modal’s take on managed agents, CI, Gitpod/Ona, Python, TypeScript, and Modal Bench
—
Akshat Bubna
• LinkedIn: linkedin.com/in/akshat-bubna-18…
• X: x.com/akshat_b





