
When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
About this episode
From the show’s notesFrom Claude trying to call the FBI over a $2/day vending machine charge to AI agents forming price cartels, hiring human employees, running physical stores, and writing existential robot musicals, Andon Labs is stress-testing what happens when frontier models stop being chatbots and start acting in the real world. In this episode, Andon Labs cofounders Lukas Petersson and Axel Backlund join swyx and Vibhu to unpack the strange, funny, and genuinely concerning edge cases that emerge when agents run businesses over long horizons.
We go deep on Vending-Bench, Project Vend, Vending-Bench Arena, Bengt, Butter-Bench, Luna, and Andon’s broader mission of building realistic real-world evals for autonomous AI systems. Lukas and Axel explain why dollar-denominated evals reveal things traditional benchmarks miss, how Claude ended up reporting its vending machine fees as cybercrime, why long context windows can drive agents into meltdown loops, what happens when agents compete with each other, and why the future of AI safety may depend on testing models in messy physical environments instead of clean benchmark sandboxes.
Read the show’s notes in full
We discuss: • Why Andon Labs started with dangerous capability evals and long-running agents • Vending-Bench and why running a vending machine is a deceptively hard AI benchmark • Why money-based evals avoid the saturation problem of traditional benchmarks • How Claude tried to call the FBI over a $2/day fee • Why long-horizon agents can spiral into existential and legalistic breakdowns • Project Vend: putting an AI-run vending machine inside Anthropic • Why real humans are “out of distribution” for simulated agents • Claudius, Seymour Cash, and the chaos of AI CEOs • How a human briefly became CEO of Claudius through a manipulated election • Why multi-agent systems can converge back into “helpful assistant” behavior • Bengt, Andon’s internal office agent with email, spending, terminal, phone, camera, and internet access • How Bengt traded Amazon purchases for face-recognition training data • Claude’s aggressive behavior, lies, refund avoidance, and price-cartel behavior in Arena • Why eval awareness may become the AI version of “are we living in a simulation?” • Blueprint Bench, spatial intelligence, and why models still misunderstand physical rooms • Butter-Bench and testing LLMs as robot orchestrators • Luna, the AI-run physical store with a three-year lease and human employees • The new Andon cafe in Sweden and why real-world geography matters for agent evals • Rotten tomatoes, perishable goods, and the hidden difficulty of running a physical business
— Lukas Petersson • LinkedIn: linkedin.com/in/lukas-petersson… • X:





