Learning Session--Cutting AI Cloud Costs
DeepRunner AI
0:00 / 0:00
Learning Session--Cutting AI Cloud Costs
30 просмотров · 6 дней назад
DeepRunner AI
14 подписчиков
30 просмотров · 6 дней назад
Cloud bills for AI/ML workloads can spiral fast. Speaker will share a case study from a previous role, where he led a cloud infrastructure overhaul for an AI-powered product.
We'll cover:
Why heavy Docker containers were silently driving up costs and slowing cold starts
How converting models to ONNX cut model size by ~70% and improved inference latency by 3x
A hands-on comparison of AWS EC2 + Auto Scaling, ECS Fargate, and Lambda for serverless inference
The event-driven architecture that solved payload size and timeout limitations
Practical lessons for running deep learning inference in production on a budget