
OpenAI’s New Agent Stack: Computer Use, Decisions API, UltraFast, Dots—Ari Weinstein & Nikunj Handa
About this episode
From the show’s notesFrom giving AI agents their own cloud computers to making Computer Use faster than humans at real-world tasks, OpenAI is pushing agents much closer to actually operating software end-to-end. In this episode, recorded immediately after OpenAI DevDay, Ari Weinstein, who leads product and engineering for Computer Use, joins swyx and Vibhu to unpack Dots, GPT-6.1, the Agents API, App Shots, and the rapid improvements that have made Computer Use dramatically faster, cheaper, and more reliable.
Ari explains why Computer Use is now “180 degrees different” from where it was months ago, how agents are learning to debug and recover from failures, why combining screenshots with accessibility data, the DOM, Playwright, and generated code changes the speed equation, and why the next frontier is making agents literally superhuman at using software.
Read the show’s notes in full
In the second half, Nikunj from OpenAI’s API team breaks down the new developer stack: async tool calling, mid-turn steering, WebSockets, UltraFast inference, the Decisions API, prompt caching, pre-warming, compaction, and the Agents API. He also shares the unusually fast story behind Decisions API, how OpenAI’s inference teams are using coding agents to optimize their own systems, and the bigger question of whether OpenAI is effectively building an “AI cloud” with higher-level primitives for agents, memory, compute, and state.
We discuss: • Why OpenAI thinks Computer Use has changed dramatically in just the last few months • Dots and what changes when every agent gets its own Linux computer • Why Computer Use can now complete some tasks faster than the average human • The path from human-level to “literally superhuman” software operation • Why modern Computer Use agents are much better at debugging and recovering from failure • How screenshots, accessibility trees, the DOM, Playwright, and generated JavaScript work together • App Shots and why they give models much richer context than ordinary screenshots • Why Computer Use can close the loop between writing software and testing it • Trust, permissions, and safety when agents can make payments or operate third-party websites • Async function calling and why models no longer need to stop reasoning while tools run • Mid-turn steering, WebSockets, and the architecture behind more responsive agents • UltraFast inference and how OpenAI is pushing frontier models toward much lower latency • The rapid internal story behind Decisions API • Why Decisions API is more than just structured outputs at low latency • GPT Live, fast tool calling, and real-time computer control • How OpenAI is using Decisions API internally for support classification and other workflows • Longer prompt caching, cache pre-warming, and building cache-aware applications • Server-side compaction vs manual compaction for long-running agent threads • What should live inside an Agents API versus inside a developer’s own harness • OpenAI as an “AI cloud” and the search for higher-level primitives beyond raw model APIs Sponsorship and business inquiries: business@latent.space





