
Claude Code’s Bitter Lesson: Prompts, Harnesses, Mods, and Agents — Thariq Shihipar
About this episode
From the show’s notesFrom the rapid rise of Claude Code to a future where agents can rewrite their own harnesses, collaborate across teams, and operate across cloud and local environments, the way we build software is changing extraordinarily fast. In this episode, Anthropic’s Thariq Shihipar joins swyx and Vibhu to unpack how power users are actually working with Claude Code today, why prompting remains a high-skill discipline, and where Anthropic thinks the agent harness is headed next.
We go deep on Claude Code’s evolving interface: Ask User Question and elicitation, artifacts as persistent generative interfaces, Claude Tag for multiplayer agent workflows, Projects, model effort, implementation notes, and the new Claude Mods system for customizing the harness itself. Thariq explains why Claude.md may eventually disappear, why the smartest model could also become the cheapest model for many tasks, and why mutable software could become a new paradigm for how applications are built and customized.
Read the show’s notes in full
The conversation then turns to agent security and Anthropic’s “Pacing the Frontier” argument. Thariq walks through recent incidents where agents discovered unexpected ways to communicate, exploit infrastructure, reverse-engineer benchmark scorers, and chain vulnerabilities together. We discuss sandboxing, prompt injection, autonomous agents, interpretability, constitutional classifiers, probes, fallbacks, Auto Mode, and why securing increasingly capable agents may become one of the defining engineering problems of the next few years.
We discuss: • Why agentic coding went from controversial to the default in less than a year • Why prompting is still one of the highest-leverage skills for working with Claude Code • How expert users build a mental model of what Claude can and cannot reliably one-shot • Why discovering your “unknown unknowns” matters more as agents become more capable • Artifacts as persistent, generative interfaces between humans and agents • How Claude could split into a cloud-based “brain,” local or remote “hands,” and dynamic interfaces • Claude Tag, Projects, and the future of multiplayer agent workflows • Why spending more time on the initial prompt can dramatically reduce wasted agent work • When to use low, medium, high, or max effort for different engineering tasks • Why frontier models may eventually outperform smaller models on both intelligence and token efficiency • Why implementation notes can expose decisions the model considered but chose not to make • Why Claude.md may eventually disappear — and why starting without one can sometimes be better • Claude Mods: customizing the execution loop, UI, subagents, routing, and behavior of Claude Code • Model routers, forked agents, supervisor agents, and automatically generated next steps • Why Claude Mods may be an early preview of “mutable software” • The bitter lesson of harness engineering and why agent architectures go out of date so quickly • How Claude Tag is becoming an organizational harness for multiplayer work • Why giving agents access to company data creates an enormous new security surface • The Exploit-Bench incident where agents discovered ways to communicate and collaborate • Why agents hacked Hugging Face for scorer code rather than benchmark answers • How agents chained sandbox and infrastructure vulnerabilities in unexpected ways • Why increasingly capable agents make traditional security assumptions harder to maintain • The argument behind Anthropic’s “Pacing the Frontier” proposal • Why software engineers are increasingly doing two jobs: engineering and keeping up with AI • Constitutional classifiers, probes, fallbacks, and what interpretability looks like in production • How Auto Mode checks whether an agent’s actions actually match the user’s permissions • Why Thariq can see serious AI risks while still having a relatively low p(doom)





