Agentic Conversations (formally mlops.community)

← Agentic Conversations (formally mlops.community)3 Aug · 32 min

Why Your AI Bill Will Double Before It Gets Better

Why Your AI Bill Will Double Before It Gets Better3 Aug32 min

<p>In this episode, we're joined by <em>Josh Collier</em>, FinOps Lead at Superhuman (formerly Grammarly), to explore what it really costs to run AI at scale and why the rules of the game changed faster than anyone expected.</p><p><br /></p><p>We discuss how AI token costs dropped 80% in two years, why that trend has sharply reversed with frontier models doubling in price, and how Josh rebuilt a single LLM workflow that cost $400k a month down to $80k by rethinking the architecture. He also shares how a cost calculator built in 15 minutes transformed the way his team estimates spend before running experiments, and why research-led optimization is the only kind that works without degrading the product.</p><p><br /></p><p>Along the way, we cover hidden costs most teams miss, the trade-off between Azure reserved capacity and OpenAI Priority Processing, why fixed subscription pricing is broken in an AI-native world, vendor lock-in risk, and what OpenAI's Guaranteed Capacity announcement really signals about where vendor relationships are heading next.</p><p><br /></p><p>Superhuman: <a href="https://superhuman.com" target="_blank" rel="ugc noopener noreferrer">https://superhuman.com</a></p><p><br /></p><p>Josh Collier: <a href="https://www.linkedin.com/in/josh-collier-945b7029/" target="_blank" rel="ugc noopener noreferrer">https://www.linkedin.com/in/josh-collier-945b7029/</a></p><p>Demetrios: <a href="https://www.linkedin.com/in/dpbrinkm" target="_blank" rel="ugc noopener noreferrer">https://www.linkedin.com/in/dpbrinkm</a></p><p><br /></p><p>Timestamps:</p><p>[00:00] OpenAI Guaranteed Capacity: what's really going on</p><p>[01:04] Josh's path into AI FinOps</p><p>[02:48] Token costs: the 80% price drop</p><p>[04:16] Why costs will only go up</p><p>[05:06] External LLMs as financial risk</p><p>[07:16] Why subscription pricing is dead</p><p>[08:22] The data residency fee nobody notices</p><p>[09:33] The cost calculator built in 15 minutes</p><p>[10:24] How it changed dev team speed</p><p>[13:00] Tracking costs by service and team</p><p>[15:33] $400k workflow rebuilt for $80k</p><p>[17:13] Why only research can optimize tokens</p><p>[20:00] Speculative decoding win</p><p>[23:11] One bad query, $40k gone</p><p>[26:00] Why Azure PTU was exhausting</p><p>[28:59] Shadow traffic load testing</p><p>[29:07] Priority processing: no brainer</p><p>[31:10] Guaranteed capacity: lock-in signal?</p><p>[32:18] The danger of multi-year AI deals</p><p>[33:28] Vendor-agnostic proxy as exit strategy</p>