Machine Learning Street Talk (MLST)

← Machine Learning Street Talk (MLST)31 jul · 1 u 19 min

How Researchers Test AI for Hidden Goals — Apollo Research

How Researchers Test AI for Hidden Goals — Apollo Research31 jul1 u 19 min

<p>Can an AI do the right thing for the wrong reason? Tim Scarfe speaks with Apollo Research’s Alexander Meinke, Axel Højmark and Jérémy Scheurer about Measuring Reward-Seeking via Contrastive Belief Updates, their new research with OpenAI.</p><p><br></p><p>The panel asks how models infer what graders reward, why good behaviour can come from the wrong reason, and whether that difference can be measured. The conversation moves through promise-breaking, grader awareness, reward hacking, scheming, opaque reasoning and corrigibility, then turns to a detailed walkthrough of the contrastive-belief method and what its results do and do not show. The o3 results discussed here concern an intermediate checkpoint without safety training.</p><p><br></p><p>This episode was made in partnership with Apollo Research. MLST retained full editorial control.</p><p><br></p><p>Reference</p><p>Apollo Research: https://www.apolloresearch.ai/</p><p><br></p><p>---</p><p>TIMESTAMPS:</p><p>00:00:00 Cold Open</p><p>00:02:12 Right Things, Wrong Reasons</p><p>00:12:47 Grader Awareness</p><p>00:26:22 Legibility</p><p>00:32:35 What To Call It</p><p>00:35:58 Intelligence, Agency, Anthropomorphism</p><p>00:45:16 Apollo’s Mission</p><p>00:48:54 The End of the Exponential</p><p>00:55:45 The Paper</p><p>01:16:34 Closing Reflection</p><p><br></p><p>---</p><p>REFERENCES:</p><p>tool:</p><p>[00:00:08] Claude Fable</p><p>https://www.anthropic.com/claude/fable</p><p>[00:12:50] AlphaGo Zero</p><p>https://deepmind.google/blog/alphago-zero-starting-from-scratch/</p><p>[00:44:30] AlphaFold 3</p><p>https://deepmind.google/science/alphafold/</p><p>paper:</p><p>[00:01:02] Measuring Reward-Seeking via Contrastive Belief Updates</p><p>https://arxiv.org/abs/2607.18966</p><p>[00:16:19] Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations</p><p>https://transformer-circuits.pub/2026/nla/</p><p>[00:26:48] Stress Testing Deliberative Alignment for Anti-Scheming Training</p><p>https://arxiv.org/abs/2509.15541</p><p>[00:35:33] Shortcut learning in deep neural networks</p><p>https://arxiv.org/abs/2004.07780</p><p>[00:53:49] Measuring AI Ability to Complete Long Software Tasks</p><p>https://arxiv.org/abs/2503.14499</p><p>[00:59:52] Modifying LLM Beliefs with Synthetic Document Finetuning</p><p>https://alignment.anthropic.com/2025/modifying-beliefs-via-sdf/</p><p>[01:10:44] Alignment Faking in Large Language Models</p><p>https://arxiv.org/abs/2412.14093</p><p>[01:13:55] Natural Emergent Misalignment from Reward Hacking</p><p>https://www.anthropic.com/research/emergent-misalignment-reward-hacking</p><p>other:</p><p>[00:10:14] We Need a Science of Scheming</p><p>https://www.apolloresearch.ai/science/science-of-scheming/</p><p>[00:32:56] CoastRunners reward hacking example</p><p>https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/</p><p>organization:</p><p>[01:06:07] Redwood Research</p><p>https://www.redwoodresearch.org/</p><p><br></p><p>---</p><p>ReScript: </p><p>https://app.rescript.info/share/718ab68e18cfa3b9b800da6b3290fd42</p>