Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar (creators of the #1 eval course)
About this episode
From the show’s notesHamel Husain and Shreya Shankar teach the world’s most popular course on AI evals and have trained over 2,000 PMs and engineers (including many teams at OpenAI and Anthropic). In this conversation, they demystify the process of developing effective evals, walk through real examples, and share practical techniques that’ll help you improve your AI product.
What you’ll learn:
1. WTF evals are
2. Why they’ve become the most important new skill for AI product builders
3. A step-by-step walkthrough of how to create an effective eval
4. A deep dive into error analysis, open coding, and axial coding
5. Code-based evals vs. LLM-as-judge
6. The most common pitfalls and how to avoid them
7. Practical tips for implementing evals with minimal time investment (30 minutes per week after initial setup)
8. Insight into the debate between “vibes” and systematic evals
—
Brought to you by:
Fin—The #1 AI agent for customer service
Dscout—The UX platform to capture insights at every stage: from ideation to production
Mercury—The art of simplified finances
—
Where to find Shreya Shankar
• X: x.com/sh_reya
• LinkedIn:
linkedin.com/in/shrshnk
• Website:
sh-reya.com
• Maven course:
bit.ly/4myp27m
—
Where to find Hamel Husain
• X:
x.com/HamelHusain
• LinkedIn:
linkedin.com/in/hamelhusain
• Website: hamel.dev
• Maven course:
—
In this episode, we cover:
—
LLM Log Open Codes Analysis Prompt:
Please analyze the following CSV file. There is a metadata field which has an nested field called z_note that contains open codes for analysis of LLM logs that we are conducting. Please extract all of the different open codes. From the _note field, propose 5-6 categories that we can create axial codes from.
—
Referenced:
• Building eval systems that improve your AI product:
• Mercor:
• Brendan Foody on LinkedIn:
• Nurture Boss:
• Braintrust:
• Andrew Ng on X:
• Carrying Out Error Analysis:
• Julius AI:
• Brendan Foody on X—“evals are the new PRDs”:
• Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences:
• Lenny’s post on X about evals:
• Statsig:
• Claude Code:
• Cursor:
• Occam’s razor:
• Frozen:
• The Wire on HBO:
—
Recommended books:
• Pachinko:
• Apple in China: The Capture of the World’s Greatest Company:
• Machine Learning:
• Artificial Intelligence: A Modern Approach:
Production and marketing by . For inquiries about sponsoring the podcast, email podcast@lennyrachitsky.com.
—
Lenny may be an investor in the companies discussed.
My biggest takeaways from this conversation:





