Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize
AI Engineer
0:00 / 0:00
Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize
17 668 просмотров · 3 месяца назад
AI Engineer
633 тыс. подписчиков
17 668 просмотров · 3 месяца назад
Most agents get tested by running a few queries and checking if it looks right. Laurie calls this the vibes problem: it doesn't catch regressions, doesn't run in CI, and doesn't tell you whether a prompt fix broke three other things. This workshop builds a complete eval pipeline from scratch on a financial analysis agent: tracing with Phoenix, reading traces before writing a single eval, categorizing failures by root cause, then building code evals, built-in LLM-as-a-judge evals, and a custom rubric with labeled examples.
The sharpest lesson: choosing the right eval matters more than tuning it. A correctness eval scored 0 out of 13 on the same agent that a faithfulness eval scored 13 out of 13, because the model doesn't know what year it is and can't verify forward-looking financial data. The workshop closes on the thing most eval content skips — experiments that let you prove a prompt change actually worked, rather than eyeballing it and calling it a win.
Speaker info:
https://x.com/seldo
/ seldo
https://github.com/seldo
Timestamps:
0:00:00 Introduction
0:00:14 Workshop Overview
0:04:31 Troubleshooting Phoenix Setup
0:05:17 Fundamentals of Evals and Tracing
0:18:44 Anatomy of an Eval Result
0:21:19 The Iteration Loop
0:26:58 Building the Financial Analysis Agent
0:33:28 Using Phoenix for Observability
0:35:38 Running Multiple Test Queries
0:38:12 Reading and Categorizing Traces
0:49:52 Implementing Code Evals
0:57:51 Built-in LLM-as-a-Judge Evals
1:03:04 Faithfulness Evaluation
1:04:35 Designing a Custom Eval Rubric
1:11:47 Running the Actionability Judge
1:19:14 Using Data Sets and Experiments
1:50:19 Final Tips and Best Practices
1:51:48 Differences Between Phoenix and Arize AX