
← LessWrong (30+ Karma)Nieuw · 19 min
“Astra’s no-CoT limits track speculative depth, not step count” by MBaert
<p> tl;dr I have tested Astra's ability to complete various long multi-step tasks without using its chain-of-thought, and found that the number of sequential steps is a poor predictor of task success. Instead, Astra's ability to solve a task seems to correlate more strongly with what I will call the speculative depth of the task. Based on experimental results, it seems less likely that Astra solves sequential no-CoT tasks purely by reasoning step-by-step in latent space. Instead, Astra appears to do some form of speculative reasoning, where intermediate results are guessed based on heuristics, and then iterated upon in parallel until they become self-consistent. This allows many multi-step tasks to be solved with far fewer serial steps than naively seems possible, especially if the initial guesses are good. Other LLMs also appear to behave like this, but to a much smaller degree.</p><p> In my previous post, I discussed how certain KV-cache sharing schemes may lead to long opaque serial paths, and introduced the LatentMathBench microbenchmark, which measures an LLM's ability to solve tasks that require many consecutive steps without using their chain-of-thought. Several other users have done more extensive no-CoT reasoning benchmarks based on more varied (and often more realistic) [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(01:59) Shortcuts</p><p>(03:41) Boolean circuits</p><p>(08:25) But how?</p><p>(10:22) Speculative reasoning</p><p>(13:01) Testing task success rate vs speculative depth</p><p>(14:38) Open questions</p> <p><i>The original text contained 3 footnotes which were omitted from this narration.</i> </p><p>---</p>
<p><b>First published:</b><br/>
September 11th, 2026 </p>
<p><b>Source:</b><br/>
<a href="https://www.lesswrong.com/posts/WFc3NkuPaYFrYuaZd/astra-s-no-cot-limits-track-speculative-depth-not-step-count?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/WFc3NkuPaYFrYuaZd/astra-s-no-cot-limits-track-speculative-depth-not-step-count</a> </p>
<p>---</p>
<p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p>
<p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1789148285/lexical_client_uploads/heoojl4sq4gdak97iuks.svg" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1789148285/lexical_client_uploads/heoojl4sq4gdak97iuks.svg" alt="Bar graph titled "LatentMathBench (bool-N) 50% success horizon (out of 50 trials)" comparing models." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1789148468/lexical_client_uploads/rpdysi8c1db4rd3clmky.svg" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1789148468/lexical_client_uploads/rpdysi8c1db4rd3clmky.svg" alt="Bar graph titled "LatentMathBench (xor-only bool-N) 50% success horizon (out of 50 trials)"." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1789148979/lexical_client_uploads/d0vvr1et7ngzpndhdomq.svg" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1789148979/lexical_client_uploads/d0vvr1et7ngzpndhdomq.svg" alt="Heatmap "LatentMathBench (bool-4) success rate for GPT-6 Astra (out of 10 trials)"." style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>