Перейти к содержимому

Scaling AI Inference: KV Cache, llm-d, and the Systems Bottleneck

Splainers HQ

0:00 / 0:00

Scaling AI Inference: KV Cache, llm-d, and the Systems Bottleneck

160 просмотров · 13 дней назад
Splainers HQ
17 подписчиков
160 просмотров · 13 дней назад
Every AI answer depends on a serving system that must move memory, schedule work, and use expensive accelerators efficiently. This explainer follows IBM researcher Danny Harnik’s SYSTOR keynote on KV-cache management, prefill/decode disaggregation, distributed GPU utilization, and the open-source llm-d project. The talk presents active systems research and engineering approaches in a fast-changing field. Hardware, model formats, and platform capabilities may evolve quickly; AI-generated narration and visuals may contain errors. ORIGINAL SOURCE Creator: Systor Conference Title: System Challenges for Scaling AI Inference Published: September 9, 2026    • System Challenges for Scaling AI Inference   ABOUT TECH-SPLAINERS Tech-Splainers provides explainers for important discussions happening in the tech world. Independent educational summary; not affiliated with or endorsed by the original creator, speaker, or IBM. Watch the original for full context. #TechSplainers #AIInfrastructure #Inference