Air Street Press

← Air Street Press30 Jun · 8 min

Compute scarcity is an engineering problem

Compute scarcity is an engineering problem30 Jun8 min

<p>Angelos Perivolaropoulos, a research engineer at ElevenLabs, on turning GPU scarcity into an inference-engineering problem: how to serve far more users on the same hardware, from batching to frontier architecture changes. Recorded at RAAIS 2026.</p><p>00:00 Introduction: ElevenLabs and the GPU squeeze</p><p>00:38 The question: how to scale when you can&#39;t add capacity</p><p>01:11 About Angelos: Scribe, speech-to-text and text-to-speech</p><p>01:56 GPU scarcity meets exponential demand</p><p>02:44 What a token actually costs: compute vs memory bandwidth</p><p>03:38 Prefill, decode and the KV cache</p><p>05:53 Batching and continuous batching (1 → 15 users/GPU)</p><p>08:37 FP8 quantization and quantize-aware training (→ 20)</p><p>11:29 Speculative decoding and multi-token prediction (→ 28)</p><p>15:13 Compressing the KV cache: TurboQuant and distillation (→ 70)</p><p>17:27 Frontier architectures: MLA, linear attention, state-space (→ 140)</p><p>20:39 Trade-offs: nothing is free</p><p>22:03 Q&amp;A: papers vs production, token subsidies, TTS evals</p>