
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
About this episode
From the show’s notesFrom routing a 200,000-token prompt across GPUs to having GLM-5.2 profile, rewrite, and optimize the kernels serving itself, inference engineering is becoming one of the most important layers of AI. In this episode, Baseten’s Philip Kiely and Ali Taha join swyx to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.
We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.
Read the show’s notes in full
The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.
We discuss: • What happens when a 200,000-token request enters an inference system • Cache-aware routing and reusing previously computed KV cache • Why prefill and decode are increasingly handled by different GPUs • When dedicated deployments become cheaper and more reliable than shared APIs • How speculative decoding uses a smaller model to accelerate a larger one • Tool calling, structured outputs, and what LLMs actually do • What it takes to support a new open model on day zero • Grafting Kimi’s vision encoder onto GLM-5.2 • Retrofitting inefficient model layers with components from other architectures • Why models sometimes collapse into repeating the same token • How hardware, kernels, and race conditions create nondeterministic failures • Preserving model fidelity while making inference faster • How quantization errors can cancel each other out • Why inference optimizations still deliver gains of 20%, 100%, and 200% • How optimized serving can make a model up to 10× faster • NVIDIA Dynamo, KV-aware routing, and distributed model serving • Speculative decoding the speculative decoder • Why local AI is about making models less dumb while data-center AI is about making them less slow • Tensor, expert, and pipeline parallelism across GPUs • Hardware-aware model design, auto-tuning, and the case against mega kernels • Rubin and why inference is becoming a systems problem • Whether modern GPUs are evolving into programmable AI ASICs • Why enormous models like Kimi K3 require GB300-class hardware • Why open-source video generation still trails Veo, Kling, and other closed models • The quadratic attention bottleneck behind long-form AI video • Autoregressive video, real-time generation, and compounding quality drift • Why future video systems may combine autoregressive and diffusion architectures • Training for inference and inference for training • Continuous post-training, deployment, evaluation, and improvement loops • How GLM-5.2 helped optimize the kernels serving GLM-5.2 • Why faster networking could unlock dramatically faster decoding • Continual learning, KV-cache compaction, and persistent model memory





