AI Post Transformers

← AI Post Transformers4 sep

KV Cache Management Faces Its First Head-to-Head Test

KV Cache Management Faces Its First Head-to-Head Test4 sep

This episode examines a comparative study of three KV cache management strategies for LLM inference — vLLM's PagedAttention memory management, H2O's static sparsification, and InfiniGen's dynamic CPU-offload selection — tested side by side on identical hardware for the first time. The standout finding: both H2O and InfiniGen hit out-of-memory errors around 10,000 tokens, less than 10% of the 128K context window modern models claim to support, revealing that many eviction-based approaches can't even survive prefill on long documents. The discussion traces why KV caches exist at all (avoiding quadratic recomputation cost), how Grouped Query Attention reduces steady-state cache size but does nothing for the transient attention-score matrix that must be materialized during prefill to decide what to evict, and why that structural gap explains the paradigms' divergent failure modes. Testing spans Llama-3.1-8B and 70B, GPT-OSS-20B, and multiple benchmark datasets across four H100 GPUs. Listeners interested in the practical limits of long-context LLM serving — and why architectural tricks like GQA don't fully solve the memory problem — will find the paper's empirical exposure of these failure points compelling.

Sources:

1. Comparative Characterization of KV Cache Management Strategies for LLM Inference — Oteo Mamo, Olga Kogiou, Hyunjin Yi, Weikuan Yu, 2026

http://arxiv.org/abs/2604.05012

2. Efficient Memory Management for Large Language Model Serving with PagedAttention — W. Kwon et al., 2023

https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention

3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Z. Zhang et al., 2023

https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models

4. InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management — W. Lee, J. Lee, J. Seo, J. Sim, 2024

https://scholar.google.com/scholar?q=InfiniGen%3A+Efficient+Generative+Inference+of+Large+Language+Models+with+Dynamic+KV+Cache+Management

5. Characterizing the Behavior and Impact of KV Caching on Transformer Inferences under Concurrency — J. Ye, J. Cernuda, A. Maurya, X.-H. Sun, A. Kougkas, B. Nicolae, 2025

https://scholar.google.com/scholar?q=Characterizing+the+Behavior+and+Impact+of+KV+Caching+on+Transformer+Inferences+under+Concurrency