
← The Growth Podcast28 Jul · 57 min
How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google
Today’s Episode
A developer posted this workflow in March, and it is the clearest picture of where PM is heading that I’ve seen all year.
Rasty Turek spent the past year building with coding agents, and he mapped how his process changed over that time. He reckons he now spends around 90% of his time on evals. His eval started as QA, and then it became the spec.
Great, now everyone agrees evals are important and will become indispensable for PMs going forward. But there is very little on how to write one.
That changes today.
I’ve now done 6 episodes on evals, and all of them start with an agent that is running and failing. So what do you do on day 0?
Daniel McKinnon was a PM on the Llama models at Meta, a boomerang who spent around 7 years there in total. He sat on Facebook’s central AI team for the entirety of its existence. He wrote enterprise evals for Gemini, Llama, and Ray-Ban Meta.
His first job at Meta was on the speech recognition team. He had to figure out how to check whether the models were any good. They weren’t called evals back then. But he’s been writing them for his entire career anyway.
In this episode you’ll learn:
* How to build an eval set from nothing
* The floor-and-ceiling method for calibrating
* How to score it and make the shipping call