
← AI Post Transformers4 days ago
Photonic Interconnects and the Race to Cut Prefill Latency
This episode examines a Lightmatter-authored paper claiming photonic interconnects can cut inference prefill latency by up to 8.5x for long-context, Mixture-of-Experts workloads. The hosts unpack why prefill has become a dominant cost center as agentic coding pushes median prompt lengths toward 96K tokens, and explain the technical distinction between compute-bound prefill and memory-bandwidth-bound decode. They dig into the physics behind the claim: copper's one-meter reach limit at 224 Gbps per lane forces multi-rack scale-out, while 3D-integrated photonics decouples I/O from a chip's shoreline, enabling far higher bandwidth density. Throughout, the co-hosts push back on taking the headline multiplier at face value, stressing that the bandwidth specs come from Lightmatter's own published sheet and the performance gains from a simulator the company itself built and controls. It's a useful listen for anyone wanting a grounded, skeptical walkthrough of interconnect physics versus vendor-reported benchmarks in AI infrastructure claims.
Sources:
1. Scaling Inference Prefill with High-Radix Photonic Interconnects — Arulselvan Madhavan, Peter Carson, Taylor Groves, Thomas Graham, 2026
http://arxiv.org/abs/2609.01821
2. Accelerating Frontier MoE Training with 3D Integrated Optics — M. Bernadskiy, P. Carson, T. Graham, T. Groves, H. J. Lee, E. Yeh, 2025
https://scholar.google.com/scholar?q=Accelerating+Frontier+MoE+Training+with+3D+Integrated+Optics
3. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized LLM Serving — Y. Zhong, S. Liu, et al., 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+LLM+Serving
4. Sarathi-Serve: Taming Throughput-Latency Tradeoff in LLM Inference — A. Agrawal, N. Kedia, A. Panwar, et al., 2024 (OSDI)
https://scholar.google.com/scholar?q=Sarathi-Serve%3A+Taming+Throughput-Latency+Tradeoff+in+LLM+Inference
5. Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast LLM Inference — Q. Li, B. Zhang, L. Ye, et al., 2024
https://scholar.google.com/scholar?q=Flash+Communication%3A+Reducing+Tensor+Parallelization+Bottleneck+for+Fast+LLM+Inference