Перейти к содержимому

Building Reliable Data Agents: Our Evaluation Framework - Fabian Marino and Felix Pitterling @Idealo

Data Berlin

0:00 / 0:00

Building Reliable Data Agents: Our Evaluation Framework - Fabian Marino and Felix Pitterling @Idealo

119 просмотров · 7 дней назад
Data Berlin
45 подписчиков
119 просмотров · 7 дней назад
To speed up internal business analytics, Idealo created a conversational AI agent capable of querying AWS Athena and their data lake directly in plain English. To guarantee accuracy without manual spot-checks, they built an automated "LLM as a judge" evaluation framework. It tests 45 distinct business scenarios across quality metrics like numeric correctness, SQL validity, and answer completeness. Integrated directly into their CI/CD workflow via GitHub Actions, every pull request automatically generates an HTML report evaluating agent performance before code is merged. Hosted by FGS Global https://fgsglobal.com/ Recording by Lukas Rieder from hibase   / lukasrieder  .