Перейти к содержимому

Deploying an AI Platform on Infrastructure You Don't Control | AI Scale Talks EP.3

Lablup Inc

0:00 / 0:00

Deploying an AI Platform on Infrastructure You Don't Control | AI Scale Talks EP.3

102 просмотра · 6 дней назад
Lablup Inc
733 подписчика
102 просмотра · 6 дней назад
The third episode of AI Scale Talks goes out to customer sites, where Backend.AI runs on infrastructure we don't control. Jonghyun Park, co-founder and CPO of Lablup, covers what we need to know about a site before installing, how we deploy on air-gapped sites with offline bundles and PyInfra, and how we support customers once the platform is in production. AI Scale Talks runs four Wednesdays under the theme From Cell to Factory. Each episode scales up one level, from a single inference engine to full AI infrastructure operations. EP.3 moves to the site level: every customer brings a different network, security policy, and hardware stack. ⏰ Chapters 00:00 Welcome and series overview 02:25 Housekeeping 02:59 Meet the speaker 04:09 Our products and customers 05:43 One stack from infrastructure to workloads 08:20 Composable deployments 09:51 Two kinds of support work 13:45 Part 1: Gathering site information 15:54 Inventory and environment files 16:18 Moving from Ansible to PyInfra 17:16 Offline bundles for air-gapped sites 17:46 Surprises during on-site installs 19:21 A TUI installer 20:15 Autopilot, an agent-based installer in testing 23:17 A test run on AWS 25:21 Sites without GPUs: mirror sites 27:25 Site Designer, in development 30:33 Part 1 recap 31:04 Part 2: Supporting customers in production 32:07 Site-specific policies and custom SDKs 33:51 General AI questions from users 35:17 Enterprise Guide: shared support knowledge 36:43 Giftbox: per-site context 39:41 Three levels of support 41:33 Part 2 recap 42:28 Final thoughts: people and agents 44:41 Wrap-up and EP.4