One Bias After Another - Daniel Fein & Max Lamparth
Safe AI Germany
0:00 / 0:00
One Bias After Another - Daniel Fein & Max Lamparth
160 просмотров · 4 месяца назад
Safe AI Germany
154 подписчика
160 просмотров · 4 месяца назад
To make AI systems helpful, we use reward models, which are essentially an automated grading system that tells the AI what a good answer looks like. But what happens when the grader itself is fundamentally biased?
On May 6th 2026, @SafeAIGermany hosted Daniel Fein & Max Lamparth from @stanford to present their latest findings on the hidden flaws inside frontier AI models (Paper: https://arxiv.org/pdf/2603.03291)
In this presentation and Q&A session, we explored how state-of-the-art AI assistants remain plagued by simple biases and their work towards solutions.
About the speakers:
Daniel is a graduate researcher in the Stanford Intelligence Systems Laboratory. His research focuses on controlling and understanding the behavior of language models. He has published work on preference learning, with broader interests in evaluation, alignment, and enabling models to reason more reliably in complex settings.
Max is a Research Fellow at the @HooverInstitution, the Stanford Intelligence Systems Laboratory, and the Stanford Center for AI Safety. His research focuses on the security and safety of language models through mechanistic interpretability, reward modeling, and robust evaluation. Before, Max was a postdoctoral fellow at Stanford and received his Ph.D. from the School of Natural Sciences at the Technical University of Munich.
About SAIGE:
Safe AI Germany (SAIGE) is building Germany's infrastructure for safer AI. Our online events are open to everyone everywhere.
🌐 safeaigermany.org
📅 Upcoming events: luma.com/saige
🐦 Follow us on LinkedIn: linkedin.com/company/safe-ai-germany
Recorded 6 May 2026.