
← AI Literacy for Leaders28 jul · 13 min
What Your AI Benchmark is Really Telling You
<p>Learn more about the host Laurence Gill at www.laurencegill.com.</p><p>In Episode 12, Laurence Gill takes that number apart. A Stanford research team called BetterBench built a 46-point audit covering benchmark design, reproducibility, and documentation, then scored 24 widely-cited tests against it. MMLU came in at 5.5. GPQA, a far less publicized test, scored double that. The reasons are specific: ambiguous question phrasing that swings scores when a comma moves, a reproducibility gap across most published benchmarks, and a quiet contamination problem where models may have already seen the answer key buried somewhere in their training data.</p><p><br /></p><p>Laurence walks through how these tests actually work, why Goodhart’s Law explains the industry’s race to game them, and how newer benchmarks like GPQA and ARC-AGI are trying to close the gap. It closes with five questions to run through before any benchmark score is allowed to inform a real decision and one open question about what happens when AI starts writing the tests that grade other AI.</p>