AI Evals Explained | How to evaluate AI Agents?
Aishwarya Srinivasan
0:00 / 0:00
AI Evals Explained | How to evaluate AI Agents?
44 432 просмотра · 2 месяца назад
Aishwarya Srinivasan
180 тыс. подписчиков
44 432 просмотра · 2 месяца назад
Most people think they’ve built a successful AI agent because it ran perfectly once in their terminal. But there’s a massive gap between a flashy demo and a robust system that works reliably in production. When your agent quietly goes sideways at 2 AM, reading transcripts one by one won’t save you.
The teams successfully shipping real-world coding agents and enterprise customer support systems aren't relying on "vibe evals" or guessing. They are actively measuring.
In this video, I break down the core mechanics of AI Agent Evaluations (Evals) in the most beginner-friendly way possible. We move past grading traditional multiple-choice machine learning models and dive into the trickier world of grading open-ended, essay-style agent workflows. By the end of this deep dive, you’ll understand how to systematically test, trace, and monitor your AI systems better than most developers building today.
Automate Your Workflow with Mistral Vibe: I use Mistral Vibe, a powerful, terminal-native AI coding agent, to fully automate building my weekly demos, readmes, and code files. Try it in your terminal or IDE here: https://mistr.al/vibe-aishwarya-yt
Chapters:
0:00 The Hardest Part of Building Agents
3:10 What is an AI Eval?
3:56 Benchmarks vs. Evals: The Critical Difference
5:46 How Evaluation Changed (Traditional ML vs. LLMs)
8:00 The Top 4 Metrics You Need to Know
9:48 Streamlining Workflows: Building My Custom Demo Generator
12:21 Three Buckets: Overlap, Semantic, and Model-Based Metrics
13:56 The Golden Dataset Rule
15:04 The 4 Ways to Grade Your AI Outputs
17:32 Setting up the Continuous Evaluation Loop
19:42 Navigating the "Whac-A-Mole" Problem
20:14 Your Tech Stack: Promptfoo, Ragas, LangSmith, and Braintrust
21:04 The One Thing to Remember
21:34 Master Agentic AI with Gen Academy
Check out my deep-dive Substack blog on AI Evals here: https://aishwaryasrinivasan.substack....
Use code EARLYBIRD for 25% off.
https://maven.com/aishwarya-srinivasa...
We are also doing a dedicated workshop on "AI for Forward Deployed Engineers"
If you are a software engineer or a solutions architect and want to pivot into becoming a forward-deployed engineer, this is a perfect workshop for you to take. We have a 50% discount going for limited time.
https://maven.com/aishwarya-srinivasa...
Subscribe for bi-weekly deep dives into production-grade AI engineering, architecture breakdowns, and practical career strategies. Drop your evaluation questions in the comments below. I read and answer as many as I can!