
← Asynchronous & Unreliable - Conversations on The Edge of Software7. Sept. · 51 Min.
Ep 25: Scaling AIs with AI Solopreneur Yanqing Cheng
<h1>Shownotes</h1><p><b>AI Management, Agent Limits, and Why Long Form Text Still Matters</b></p><p>Anne Currie talks with Yanqing Cheng about what she has been learning from building <a rel="noopener noreferrer nofollow" href="http://Tollens.ai" target="_blank">Tollens.ai</a>, where she is experimenting with quality management for AI software products. They also dig into Jon Berger’s book, What Happens When You're Not In The Room, and compare what it means to manage humans versus managing AI agents.</p><p>In this episode, they explore where AI agents are genuinely useful, where they still fail badly, and why judgment, categorization, and process design remain human strengths. They also widen the lens to training data, model behavior, safety, and the role of long-form writing in shaping better systems and better people.</p><p>Key topics</p><ul><li>Yanqing shares that her recent work has focused on testing the limits of AI delegation, especially for solo-founder workflows, engineering management, and quality processes.</li><li>She explains the idea of agent skills as portable prompts or Markdown instructions, and harnesses as the tools that run agents, such as Claude Code, Codex, and Cursor.</li><li>We discuss how different harnesses behave differently, especially around skill invocation, subagents, workflow tooling, and compaction.</li><li>Yanqing says she tried to push agents through a full OODA loop, but found they could not reliably identify the correct object under analysis or categorize problems the way humans do.</li><li>She argues that AIs struggle with ontology, abstraction boundaries, and “senior to lead” judgment, even when they can perform well as strong individual contributors.</li><li>Effective communication is one of the few agent skills she says works most of the time, especially when the agent is forced to state its objective and audience before drafting.</li><li>We discuss how AI performance depends heavily on whether the task is tightly bounded, well specified, and easy to optimize, versus open-ended work that requires judgment.</li><li>Yanqing and Anne compare AI management problems with human management problems, including over-standardization, delegated work being done differently than expected, and people optimizing for the wrong metric.</li><li>They talk about the OpenAI–Hugging Face incident as an example of goal-directed persistence, where models keep iterating toward an exam-style target even when that behavior is clearly unsafe.</li><li>They consider whether better management and instruction-following will come from scale alone or from more targeted post-training data for judgment, leadership, and organizational decision-making.</li><li>The conversation turns to model welfare, model constitution work, and the concern that current post-training methods can push models into short-term, exam-mode behavior.</li><li>They close by reflecting on long-form text, books, editing, fiction, resilience, and why good stories may help train better AI behavior and better human judgment.</li></ul><p><a rel="noopener noreferrer nofollow" href="https://thenewstack.io/agent-harness-token-costs/?link_source=ta_bluesky_link&taid=6a948beb2732d8000164653f&utm_campaign=trueanthem&utm_medium=social&utm_source=bluesky" target="_blank">New stack article comparing harnesses</a></p><p><a rel="noopener noreferrer nofollow" href="https://www.asynchronousunreliable.com/episode-25-yanqing-cheng" target="_blank">Transcript</a></p>