
Devin’s 80% Moment: Background Agents, 7x PRs, & End of Hand-Held Coding — Walden Yan & Cole Murray
About this episode
From the show’s notesFrom coining “context engineering” to building the infrastructure behind Devin’s 7x PR growth and jump from 16% to 80% of commits across Cognition repos, Walden Yan has had a front-row seat to the background-agent shift. In this episode, Cognition co-founder and CPO Walden Yan joins swyx alongside Cole Murray, creator of OpenInspect, to unpack why everyone is building their own Devin, what changed after the December 2025 model inflection, and why “spec to pull request” is now becoming a real production workflow.
We go deep on the architecture of background agents: harness-in-the-box vs out-of-the-box, why Devin separates the “brain” from the machine, why repo setup is still one of the hardest problems, why Docker is not always enough, and how full VMs, snapshots, scoped secrets, GitHub bots, Slack integrations, and video-based testing all fit together. Walden and Cole also dig into memory, MCP limitations, multi-agent orchestration, AI code review, SRE auto-triage, PMs shipping code from Slack, Windsurf 2.0, hybrid frontier/sub-frontier systems, and the real failure mode of uncontrolled vibe coding: your codebase regressing to your worst engineer.
Read the show’s notes in full
We discuss: • Why the engineering world is waking up to background agents and cloud agents • The December 2025 model inflection that made spec-to-PR workflows practical • Devin’s 7x merged PR growth and rise from 16% to 80% of commits • Why Cole built OpenInspect as an open-source background-agent system • Why background agents may become critical internal infrastructure • The economics of $20/seat agent products and why monetization is tricky • What Cognition actually sells beyond Devin: infra, onboarding, integrations, and adoption • Harness in the box vs out of the box, and why architecture matters • Why Devin separates the brain from the machine for security and permissions • Repo setup, scoped secrets, Docker Compose, and agent-ready dev environments • Why full VMs matter when agents need to run real applications and test them • Android, macOS, Windows, nested virtualization, and machine-specific agent work • Why testing is much harder than “computer use” • Screenshots, video verification, and the “I know it works” merge moment • GitHub UX, Devin Review, AI reviewers, and agents responding to PR comments • Why MCP alone is not enough for first-class Slack and enterprise integrations • Memory, Knowledge, skills, Claude.md, and why retrieval is still unsolved • Devin’s auto-generated memories and the challenge of memory pruning • Always-on agents as permanent PMs for issues, tickets, and product areas • Sub-agents, meta-Devin management, and what multi-agent systems actually add • Why pure auto-merge vibe coding breaks down after about two weeks • AI code smells, lint rules, reward hacking, and Semgrep for agent-written code • GitAI, inline context, and preserving the “why” behind code changes • Local testing, mock servers, older codebases, and preparing companies for agents • Windsurf 2.0 and the handoff between local foreground agents and cloud background agents • SRE auto-triage, support workflows, and agents as first responders • PMs, marketing, and non-engineers creating pull requests from Slack • AI agent budgets, $1k–$5k per engineer spend, and hybrid frontier/sub-frontier systems • The rise of autonomous coding factories and who Cognition is hiring





