Heavybit Podcasts

← Heavybit Podcasts20 aug · 53 min

Ep. #4, The New Big Data of Inference with Junchen Jiang

Ep. #4, The New Big Data of Inference with Junchen Jiang20 aug53 min

On episode 4 of Lab Notes, Amir Zohrenejad speaks with Junchen Jiang about why KVCache may be better understood as reusable, AI-native data rather than a temporary inference optimization. They explore how LMCache and CacheBlend can reduce redundant computation, move context across distributed inference systems, and help support increasingly complex AI agents. The conversation also covers multimodal workloads, open-source infrastructure, and the future of AI systems research.