

Why Medical AI Needs a Referee | Protege's Engy Ziedan
About this episode
From the show’s notesDaisy Wolf and Eva Steinman are joined by Engy Ziedan, co-founder and Chief Scientific Officer of Protege, to discuss why medical AI has a measurement problem, and why scoring well on a benchmark doesn't necessarily mean a model is ready for the hospital.
Engy explains why healthcare AI needs independent evaluations that go beyond static exams and measure how models actually perform in real-world clinical workflows. They explore the risks of subtle bias and misalignment, why the same model can rank differently depending on how it's prompted or tested, and what happens as AI becomes more personalized and changes faster than traditional healthcare quality systems can keep up.
Read the show’s notes in full
The conversation also gets into Protege's role as an independent evaluator, how contaminated training data can undermine benchmarks, and why the future of medical AI may require continuous monitoring rather than occasional testing.
Resources:
Read our insights piece: a16z.news/p/the-oracle-problem-…
Follow Engy Ziedan on X: x.com/engyziedan
Follow Daisy Wolf on X: x.com/daisydwolf
Follow Eva Steinman on X: x.com/evajsteinman
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: twitter.com/eriktorenberg





