Перейти к содержимому

Learning Session--Cutting AI Cloud Costs

DeepRunner AI

0:00 / 0:00

Learning Session--Cutting AI Cloud Costs

30 просмотров · 6 дней назад
DeepRunner AI
14 подписчиков
30 просмотров · 6 дней назад
Cloud bills for AI/ML workloads can spiral fast. Speaker will share a case study from a previous role, where he led a cloud infrastructure overhaul for an AI-powered product. We'll cover: Why heavy Docker containers were silently driving up costs and slowing cold starts How converting models to ONNX cut model size by ~70% and improved inference latency by 3x A hands-on comparison of AWS EC2 + Auto Scaling, ECS Fargate, and Lambda for serverless inference The event-driven architecture that solved payload size and timeout limitations Practical lessons for running deep learning inference in production on a budget