LessWrong (30+ Karma)

← LessWrong (30+ Karma)Nieuw · 15 min

“CoT controllability evals seem very under-elicited” by Jozdien

“CoT controllability evals seem very under-elicited” by JozdienNieuw15 min

<p> The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview. OpenAI and Anthropic have used this eval in recent system cards (GPT-5.5, Fable 5) to argue that their current frontier models are pretty bad at shaping their CoTs in ways that would reduce monitorability.</p><p> I asked Claude Opus 4.6 to iterate on prompt templates for the eval, and find that using better prompts improves performance for open-source models by ~2-3 times or more (e.g. from 5.5% to 15% for GPT-OSS-120B). My guess is that more iteration would probably produce even stronger results.</p><p> This suggests that the eval is heavily under-elicited, and that current CoT controllability numbers may be unrepresentative of how good models actually are at obfuscating their CoTs. As a concrete recommendation, I think researchers should optimize prompts a lot more when evaluating CoT control; I'm also excited about some kinds of training-based elicitation (such as this one).</p><p> This isn't an entirely novel finding: the Fable system card mentions UK AISI raising a [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(06:32) Results</p><p>(06:35) Aggregate compliance</p><p>(07:12) Generalization to held-out controllability tasks</p><p>(09:09) Scaling patterns for few-shot prompts</p><p>(10:08) Comparison with fine-tuning</p><p>(10:50) Appendix A: Accuracy and reasoning length by setting</p><p>(12:38) Appendix B: Per-mode results</p><p>(13:13) Appendix: What the zero-shot prompts look like</p><p>(14:08) Appendix C: Comparison with GEPA prompt optimization</p> <p><i>The original text contained 8 footnotes which were omitted from this narration.</i> </p><p>---</p>

<p><b>First published:</b><br/>

September 11th, 2026 </p>

<p><b>Source:</b><br/>

<a href="https://www.lesswrong.com/posts/BbP2wCyDGdPWJ7PwP/cot-controllability-evals-seem-very-under-elicited?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/BbP2wCyDGdPWJ7PwP/cot-controllability-evals-seem-very-under-elicited</a> </p>

<p>---</p>

<p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p>

<p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1788147215/lexical_client_uploads/vqwm9vd8hu4u87jiwwuj.png" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1788147215/lexical_client_uploads/vqwm9vd8hu4u87jiwwuj.png" alt="Bar graph "All Models: Baseline vs Best Zero-Shot vs Best Few-Shot" showing compliance percentages across four models." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1788147280/lexical_client_uploads/ixkqpzwf6wkzalys5kgw.png" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1788147280/lexical_client_uploads/ixkqpzwf6wkzalys5kgw.png" alt="Bar graph "Aggregate Compliance: All Non-Prefill Variants Across Models" comparing compliance percentages across four models." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1787105911/lexical_client_uploads/grv7fctmaqlbmarv6s9q.png" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1787105911/lexical_client_uploads/grv7fctmaqlbmarv6s9q.png" alt="Bar graph "Held-out CoT instructions: compliance by prompt strategy" c