High-Performance LLM/VLM Inference with MLxcel | AI Scale Talks EP.1
Lablup Inc
0:00 / 0:00
High-Performance LLM/VLM Inference with MLxcel | AI Scale Talks EP.1
270 просмотров · 13 дней назад
Lablup Inc
733 подписчика
270 просмотров · 13 дней назад
Lablup's first global webinar series opens with MLxcel, our open source inference engine for LLM and VLM serving on unified memory hardware. Jeongkyu Shin, CEO of Lablup, covers why local inference hits a memory wall, how MLxcel is built, and what it takes to serve models across Apple Silicon and NVIDIA GPUs.
AI Scale Talks runs four Wednesdays under the theme From Cell to Factory. Each episode scales up one level, from a single inference engine to full AI infrastructure operations. EP.1 is the microscope end: one machine, one engine, and the memory bandwidth that decides everything.
⏰ Chapters
00:00 Welcome and series overview
02:04 Housekeeping
03:27 Introducing MLxcel
04:45 Memory, not compute, is the bottleneck
05:39 Unified Memory Architecture
06:55 Existing engines and their limits
10:19 What MLxcel covers today
11:46 From a research tool to a production engine
15:30 Architecture: control plane and core
18:14 Model Surgery
22:06 Continuous Batching
24:02 Embedding models
24:27 Benchmarks and memory-bound inference
26:56 Distributed inference across machines
29:12 Install and metrics
30:05 Releases and roadmap
33:31 Live demo
40:39 Q&A
MLxcel is open source under Apache 2.0: https://github.com/lablup/mlxcel
Lablup builds Backend.AI, the operationtes AI workloads across a broad range of accelerators and at any scale.
https://www.lablup.com | https://backend.ai