
Disaggregating Prefill, Decode, Attention, and FFN for Agentic Inference
This episode examines when it actually pays to split LLM inference hardware into four specialized pools rather than the now-standard two-way prefill/decode split, based on a paper proposing PDAF (prefill-attention, prefill-FFN, decode-attention, decode-FFN). It traces the reasoning from why agentic workloads — which can hit context sizes of 100,000+ tokens through repeated tool calls — strain hardware differently than chatbot traffic, through the compute-bound nature of prefill versus the memory-bandwidth-bound nature of decode, and why attention and FFN sublayers batch so differently that combining them on one device forces a similar compromise. It covers prior production systems (DistServe, Splitwise, StepFun's Step-3) that motivated these splits, and introduces the authors' HeteroPanacea simulator, validated against a real 8-node NVIDIA B200 cluster, which they use to search for when the reported up to 2.06x throughput gain actually materializes versus when the added complexity isn't worth it. Listeners interested in LLM serving infrastructure will find a grounded, skeptical take on a systems paper that resists overselling its own headline number.
Sources:
1. When Does Disaggregation Pay? Simulating Prefill--Decode--Attention--FFN Specialization for Agentic LLM Inference — Przemyslaw Forys, Haoran Wu, Can Xiao, Jiayi Nie, Tony Liu, Rika Antonova, Timothy Jones, Robert Mullins, Wayne Luk, Aaron Zhao, George A. Constantinides, 2026
http://arxiv.org/abs/2608.03741
2. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving — Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, H. Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+Large+Language+Model+Serving
3. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — P. Patel, E. Choukse, C. Zhang, A. Shah, I. Goiri, S. Maleki, R. Bianchini, 2024
https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting
4. Step-3 Is Large yet Affordable: Model-System Co-Design for Cost-Effective Decoding — StepFun, 2025
https://scholar.google.com/scholar?q=Step-3+Is+Large+yet+Affordable%3A+Model-System+Co-Design+for+Cost-Effective+Decoding
5. Mooncake: Trading More Storage for Less Computation — A KVCache-Centric Architecture for Serving LLM Chatbot — R. Qin, Z. Li, W. He, J. Cui, F. Ren, M. Zhang, Y. Wu, W. Zheng, X. Xu, 2025
https://scholar.google.com/scholar?q=Mooncake%3A+Trading+More+Storage+for+Less+Computation+%E2%80%94+A+KVCache-Centric+Architecture+for+Serving+LLM+Chatbot