LessWrong (30+ Karma)

← LessWrong (30+ Karma)4 dagen geleden · 14 min

“Yet another concerning result on Astra’s no-CoT capabilities” by Christine Corry

“Yet another concerning result on Astra’s no-CoT capabilities” by Christine Corry4 dagen geleden14 min

<p> This is a research update for an on-going replication of no-CoT evals done as part of the Second Look Fellowship. In following posts, we will run more comprehensive replications of previous work and release open source tooling for no-CoT eval elicitation. Code can be found here.</p><p><strong> tl;dr</strong></p><ul> <li value="1">We replicate experiments from Greenblatt 2025 and Greenblatt 2026 on GPT-6-Astra, on the same items and protocol as our previous update on Fable 5, Opus 5, Opus 4.5, and GPT-5.6-Sol, plus Gemini 3.1 Pro, Kimi k3, and Fable 5.1.</li><li value="2">We find that Astra is a qualitative jump in no-CoT capabilities over all datasets. <ul> <li value="1">4-hop questions: Astra achieves 31% at baseline, where every other model tested scores at 1-3%</li><li value="2">3-hop questions: 70% against previous best of 22% (Gemini 3.1 Pro)</li><li value="3">Neel Nanda and Rohan Subramani report the same jump independently. Our work qualitatively replicates these results.</li></ul></li><li value="3">Astra sees more uplift from filler tokens and repeats than previous models<ul> <li value="1">4-hop performance is doubled from baseline (31%) to peak filler condition (63% at )</li><li value="2">3-hop accuracy jumps from 70% to 85%<ul> <li value="1">Filler tokens and problem repeats raise accuracy monotonically across the full range we tested</li></ul></li><li value="3">Dylan Xu, SebastianP, &amp; Alek [...]</li></ul></li></ul> <p>---</p><p><strong>Outline:</strong></p><p>(00:29) tl;dr</p><p>(02:54) Background</p><p>(03:51) Previous work</p><p>(04:49) Datasets</p><p>(06:17) Evaluation design</p><p>(07:41) Eliciting no-CoT</p><p>(08:00) Results</p><p>(08:03) 4-Hop</p><p>(08:31) Utilization of filler tokens / problem repeats</p><p>(10:23) Per-dataset results</p><p>(10:48) Per-model profiles</p><p>(11:00) Discussion</p><p>(12:59) Related work</p> <p><i>The original text contained 6 footnotes which were omitted from this narration.</i> </p><p>---</p>

<p><b>First published:</b><br/>

September 13th, 2026 </p>

<p><b>Source:</b><br/>

<a href="https://www.lesswrong.com/posts/tz5WvDouXKbiWJG8B/yet-another-concerning-result-on-astra-s-no-cot-capabilities?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/tz5WvDouXKbiWJG8B/yet-another-concerning-result-on-astra-s-no-cot-capabilities</a> </p>

<p>---</p>

<p>Narrated b