
← Max Agency13 aug · 1 u 08 min
How Unify cut its AI agent costs 95% in two weeks
<p>Connor Heggie spent his early career on a fifteen-person self-driving startup run like a research lab, then moved to Scale AI's mapping team, before becoming the co-founder and CTO at Unify. Unify builds agents for go-to-market teams, and for the last two years, one of AI's big mainstream narratives has been automating away the sales rep entirely. Connor and his team built the opposite: an agent that gives every sales rep "an engineer in their back pocket." He walks through how Unify's harness evolved from million-agent batch jobs to a chat product, and unpacks the engineering that makes it cost-effective to run at scale.</p><p>–</p><p>We also discuss:</p><ul><li>How Unify cut 90-95% of costs two weeks before launch</li><li>The 15-requests-per-second ceiling inside OpenAI's prompt cache</li><li>Why your LLM judge must be a different model family</li><li>The Speed Audit: why one at a time beats a table of 1,000</li><li>Why Unify's subagents are just a function call</li><li>What working on self-driving taught Connor about running evals</li></ul><p>–</p><p>Timestamps:<br>00:00 Introduction</p><p>01:30 "Go-to-market is a search problem"</p><p>06:10 The old workflow: drag-and-drop nodes over a million-row table</p><p>09:40 Why the harness is similar to a coding agent's</p><p>10:50 Running durable agents in the cloud without a full VM</p><p>13:15 Giving models pandas-like superpowers over a live table</p><p>15:50 Why Unify's subagents are just a function call</p><p>18:25 The 15-requests-per-second limit hiding inside OpenAI's cache</p><p>24:20 Optimizing for prompt caching hit rates</p><p>28:20 Fork versus child subagents</p><p>32:35 The Speed Audit: why one at a time beats a table of 1,000</p><p>37:50 Locking memory to keys instead of letting the agent freestyle</p><p>44:15 Why AI isn’t taking over sales</p><p>49:00 What working on self-driving taught Connor about running evals</p><p>52:46 Why your LLM judge must be a different model family</p><p>54:45 Ditching full VMs for Monty, a Python REPL that suspends</p><p>58:55 Semantic merge sort: why Connor is obsessed with RLMs</p><p>59:53 How Unify cut 90-95% of costs two weeks before launch</p><p>1:01:15 The case for "semantic linters" over skill files</p><p>1:04:55 Why Unify runs mostly on OpenAI, their first investor</p><p>1:06:12 Why 10x cheaper tokens still lose on tool efficiency</p><p>1:07:24 Why open-source models don’t make economic sense (yet)</p><p>–</p><p>Referenced:</p><ul><li><a href="https://www.anthropic.com/">Anthropic</a></li><li><a href="https://chatgpt.com/">ChatGPT</a></li><li><a href="https://claude.ai/">Claude</a></li><li><a href="https://www.anthropic.com/claude/fable">Claude Fable 5</a></li><li><a href="https://www.anthropic.com/news/claude-sonnet-5">Claude Sonnet 5</a></li><li><a href="https://www.langchain.com/blog/introducing-context-hub">Context Hub</a></li><li><a href="https://www.langchain.com/deep-agents">Deep Agents</a></li><li><a href="https://huggingface.co/zai-org/GLM-5.2">GLM-5.2</a></li><li><a href="https://openai.com/index/introducing-gpt-5-4/">GPT-5.4</a></li><li><a href="http://Helm.ai/">Helm.ai</a></li><li><a href="https://smith.langchain.com/">LangSmith</a></li><li><a href="https://openai.com/">OpenAI</a></li><li><a href="https://bellard.org/quickjs/">QuickJS</a></li><li><a href="https://arxiv.org/abs/2512.24601">Recursive Language Models (RLM)</a></li><li><a href="https://www.salesforce.com/">Salesforce</a></li><li><a href="https://scale.com/">Scale AI</a></li><li><a href="https://anthropic.com/engineering/managed-agents">Scaling Managed Agents</a></li><li><a href="https://www.unifygtm.com/">Unify</a></li></ul><p>–</p><p>Where to find Connor:</p><ul><li><a href="https://x.com/HeggieConnor">Twitter/X</a></li><li><a href="https://www.linkedin.com/in/connor-heggie/">LinkedIn</a></li></ul><p>–</p><p>Where to find Harrison:</p><ul><li><a href="https://x.com/hwchase17">Twitter/X</a></