Перейти к содержимому

Less LLM Hallucinations: Merlin-Arthur Protocols Explained

AI Coffee Break with Letitia

0:00 / 0:00

Less LLM Hallucinations: Merlin-Arthur Protocols Explained

9 798 просмотров · 2 дня назад
AI Coffee Break with Letitia
65 тыс. подписчиков
9 798 просмотров · 2 дня назад
(Advertisement/ Dauerwerbesendung) LLMs love to hallucinate: If you give a model a document and it answers correctly, how do you know the answer actually came from that document, and not from pre-training, guessing, or some spurious clue? This video is about work we did at Aleph Alpha Research on Merlin-Arthur protocols for grounding language models. The setup turns question answering into a game between Merlin, Morgana, and Arthur: Merlin keeps the evidence Arthur needs, Morgana removes it, and Arthur has to learn when the available evidence is enough to answer and when it should abstain. We go through why this is still a problem in retrieval-augmented generation (RAG), why citations alone do not solve it, how AtMan is used to identify relevant evidence, and why Merlin and Morgana co-evolve with Arthur during training. We also get to the information-theoretic part: completeness and soundness give us a lower bound on mutual information, which we turn into a grounding score with the Explained Information Fraction. Then, we look at experiment numbers. All this is because we want to know not just whether the model is correct, but whether it is correct because of the information we gave it. 📃 Deiseroth, B., Höth, M. H., Kersting, K., & Parcalabescu, L. (2025). Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols. https://arxiv.org/abs/2512.11614 📃 AtMan paper: https://openreview.net/forum?id=PBpEb... Blog post: https://aleph-alpha.com/en/blog/bound... Outline: 00:00 Being probably correct is not enough 01:48 The problems with RAG 03:45 What about citations? 04:25 Going against hallucinations 04:52 Arthur, Merlin and Morgana 07:20 Merlin and Morgana implementation 09:45 Training with Merlin and Morgana 11:54 On-the-fly data augmentation 13:39 Co-evolution: Cow on the beach example 16:44 Groundedness Measured Properly 22:19 Results 27:20 LLMs love to guess 28:40 Limitations 29:59 Takeaways AI Coffee Break Merch! 🛍️ https://ai-coffee-break-with-letitia.... Thanks to our Patrons who support us in Tier 2, 3, 4: 🙏 Vignesh Valliappan, Ivan Janov, Sunny Dhiana, Andy Ma ▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀ 🔥 Optionally, pay us a coffee to help with our Coffee Bean production! ☕ Patreon:   / aicoffeebreak   Ko-fi: https://ko-fi.com/aicoffeebreak Join this channel as a Bean Member to get access to perks:    / @aicoffeebreak   ▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀ 🔗 Links: AICoffeeBreakQuiz:    / aicoffeebreak   Twitter / X:   / aicoffeebreak   LinkedIn:   / letitia-parcalabescu   Threads: https://www.threads.net/@ai.coffee.break Bluesky: https://bsky.app/profile/aicoffeebrea... Reddit:   / aicoffeebreak   YouTube:    / aicoffeebreak   Substack: https://aicoffeebreakwl.substack.com/ Web: https://explanationmark.de/letitia https://aicoffeebreak.com #AICoffeeBreak #MsCoffeeBean #MachineLearning #AI #research​ Video editing: Nils Trost Music 🎵 : Space Navigator – Sarah, The Illstrumentarist