Scaling AI Inference: KV Cache, llm-d, and the Systems Bottleneck
Splainers HQ
0:00 / 0:00
Scaling AI Inference: KV Cache, llm-d, and the Systems Bottleneck
160 просмотров · 13 дней назад
Splainers HQ
17 подписчиков
160 просмотров · 13 дней назад
Every AI answer depends on a serving system that must move memory, schedule work, and use expensive accelerators efficiently. This explainer follows IBM researcher Danny Harnik’s SYSTOR keynote on KV-cache management, prefill/decode disaggregation, distributed GPU utilization, and the open-source llm-d project.
The talk presents active systems research and engineering approaches in a fast-changing field. Hardware, model formats, and platform capabilities may evolve quickly; AI-generated narration and visuals may contain errors.
ORIGINAL SOURCE
Creator: Systor Conference
Title: System Challenges for Scaling AI Inference
Published: September 9, 2026
• System Challenges for Scaling AI Inference
ABOUT TECH-SPLAINERS
Tech-Splainers provides explainers for important discussions happening in the tech world.
Independent educational summary; not affiliated with or endorsed by the original creator, speaker, or IBM. Watch the original for full context.
#TechSplainers #AIInfrastructure #Inference