Building Reliable Data Agents: Our Evaluation Framework - Fabian Marino and Felix Pitterling @Idealo
Data Berlin
0:00 / 0:00
Building Reliable Data Agents: Our Evaluation Framework - Fabian Marino and Felix Pitterling @Idealo
119 просмотров · 7 дней назад
Data Berlin
45 подписчиков
119 просмотров · 7 дней назад
To speed up internal business analytics, Idealo created a conversational AI agent capable of querying AWS Athena and their data lake directly in plain English.
To guarantee accuracy without manual spot-checks, they built an automated "LLM as a judge" evaluation framework. It tests 45 distinct business scenarios across quality metrics like numeric correctness, SQL validity, and answer completeness.
Integrated directly into their CI/CD workflow via GitHub Actions, every pull request automatically generates an HTML report evaluating agent performance before code is merged.
Hosted by FGS Global https://fgsglobal.com/
Recording by Lukas Rieder from hibase / lukasrieder .