Neural intel Pod

← Neural intel Pod26 Jul · 36 min

Why "Hidden Reasoning" in Filler Tokens Changes AI Safety Forever

Why "Hidden Reasoning" in Filler Tokens Changes AI Safety Forever26 Jul36 min

<p>Welcome back to <strong>Neural Intel</strong>. Today we’re diving into the Mechanistic Interpretability research (Brauer et al., 2026) that proves frontier-scale models are decoupling their internal computation from surface-level tokens.We analyze how <strong>DeepSeek V3</strong> and <strong>Kimi K2</strong> utilize filler tokens as a computational substrate to improve accuracy on multi-hop tasks, such as 2-fact addition and complex systems of equations. We go beyond the abstract to discuss:</p><ul><ul><li><strong>The Mechanistic Relay:</strong> How attention shifts from the question to a question-filler-answer relay.</li></ul><ul><li><strong>Causal Evidence:</strong> How <strong>KV-cache transplants</strong> proved that information held in the filler tokens—not just the final position—causally drives the model&#39;s answer.</li></ul><ul><li><strong>Unsupervised Decoding:</strong> The four-stage pipeline that uses the <strong>logit lens</strong>, cross-example mean subtraction, and LLM judges to read the residual stream without ground-truth labels.</li></ul><p>This episode is essential for <strong>The Architect</strong> and <strong>The Researcher</strong> looking to understand why Chain-of-Thought (CoT) monitorability is a &quot;fragile safety property&quot; and how we can close the gap using interpretability.</p><p><strong>Join the Conversation:</strong></p></ul><ul><ul><li><strong>X/Twitter:</strong> <a href="https://www.google.com/url?sa=E&q=https%3A%2F%2Fx.com%2Fneuralintelorg" target="_blank" rel="noopener noreferrer">@neuralintelorg</a></li></ul><ul><li><strong>Deep Dives:</strong> <a href="https://www.google.com/url?sa=E&q=https%3A%2F%2Fneuralintel.org" target="_blank" rel="noopener noreferrer">neuralintel.org</a></li></ul></ul><p><br></p>