Deploying an AI Platform on Infrastructure You Don't Control | AI Scale Talks EP.3
Lablup Inc
0:00 / 0:00
Deploying an AI Platform on Infrastructure You Don't Control | AI Scale Talks EP.3
102 просмотра · 6 дней назад
Lablup Inc
733 подписчика
102 просмотра · 6 дней назад
The third episode of AI Scale Talks goes out to customer sites, where Backend.AI runs on infrastructure we don't control. Jonghyun Park, co-founder and CPO of Lablup, covers what we need to know about a site before installing, how we deploy on air-gapped sites with offline bundles and PyInfra, and how we support customers once the platform is in production.
AI Scale Talks runs four Wednesdays under the theme From Cell to Factory. Each episode scales up one level, from a single inference engine to full AI infrastructure operations. EP.3 moves to the site level: every customer brings a different network, security policy, and hardware stack.
⏰ Chapters
00:00 Welcome and series overview
02:25 Housekeeping
02:59 Meet the speaker
04:09 Our products and customers
05:43 One stack from infrastructure to workloads
08:20 Composable deployments
09:51 Two kinds of support work
13:45 Part 1: Gathering site information
15:54 Inventory and environment files
16:18 Moving from Ansible to PyInfra
17:16 Offline bundles for air-gapped sites
17:46 Surprises during on-site installs
19:21 A TUI installer
20:15 Autopilot, an agent-based installer in testing
23:17 A test run on AWS
25:21 Sites without GPUs: mirror sites
27:25 Site Designer, in development
30:33 Part 1 recap
31:04 Part 2: Supporting customers in production
32:07 Site-specific policies and custom SDKs
33:51 General AI questions from users
35:17 Enterprise Guide: shared support knowledge
36:43 Giftbox: per-site context
39:41 Three levels of support
41:33 Part 2 recap
42:28 Final thoughts: people and agents
44:41 Wrap-up and EP.4