Перейти к содержимому

The GPU That Wasn't Big Enough: Serving LLMs at Scale with Mumshad Mannambeth

LIBREMINDS

0:00 / 0:00

The GPU That Wasn't Big Enough: Serving LLMs at Scale with Mumshad Mannambeth

279 просмотров · 8 дней назад
LIBREMINDS
73 подписчика
279 просмотров · 8 дней назад
A beginner-friendly deep dive into GPUs and the challenges of serving Large Language Models at scale, presented at Cloud Native Summit Kerala 2026. Mumshad Mannambeth from KodeKloud breaks down the fundamentals: what LLM models really are (billions of numbers in a file), why CPUs can't handle them (limited parallel processing), and why GPUs with VRAM are essential for AI workloads. He explains the three critical metrics for LLM infrastructure: computing power (TFLOPS), VRAM size, and memory bandwidth. From simple house price predictions to transformer models with billions of parameters, this talk demystifies the hardware requirements behind ChatGPT and similar applications. Perfect for cloud-native engineers looking to understand AI infrastructure basics. Links: https://cnskerala.in #LLM #GPU #CloudNative #KodeKloud #AIInfrastructure