Less LLM Hallucinations: Merlin-Arthur Protocols Explained
AI Coffee Break with Letitia
0:00 / 0:00
Less LLM Hallucinations: Merlin-Arthur Protocols Explained
9 798 просмотров · 2 дня назад
AI Coffee Break with Letitia
65 тыс. подписчиков
9 798 просмотров · 2 дня назад
(Advertisement/ Dauerwerbesendung) LLMs love to hallucinate: If you give a model a document and it answers correctly, how do you know the answer actually came from that document, and not from pre-training, guessing, or some spurious clue?
This video is about work we did at Aleph Alpha Research on Merlin-Arthur protocols for grounding language models. The setup turns question answering into a game between Merlin, Morgana, and Arthur: Merlin keeps the evidence Arthur needs, Morgana removes it, and Arthur has to learn when the available evidence is enough to answer and when it should abstain.
We go through why this is still a problem in retrieval-augmented generation (RAG), why citations alone do not solve it, how AtMan is used to identify relevant evidence, and why Merlin and Morgana co-evolve with Arthur during training. We also get to the information-theoretic part: completeness and soundness give us a lower bound on mutual information, which we turn into a grounding score with the Explained Information Fraction. Then, we look at experiment numbers.
All this is because we want to know not just whether the model is correct, but whether it is correct because of the information we gave it.
📃 Deiseroth, B., Höth, M. H., Kersting, K., & Parcalabescu, L. (2025). Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols. https://arxiv.org/abs/2512.11614
📃 AtMan paper: https://openreview.net/forum?id=PBpEb...
Blog post: https://aleph-alpha.com/en/blog/bound...
Outline:
00:00 Being probably correct is not enough
01:48 The problems with RAG
03:45 What about citations?
04:25 Going against hallucinations
04:52 Arthur, Merlin and Morgana
07:20 Merlin and Morgana implementation
09:45 Training with Merlin and Morgana
11:54 On-the-fly data augmentation
13:39 Co-evolution: Cow on the beach example
16:44 Groundedness Measured Properly
22:19 Results
27:20 LLMs love to guess
28:40 Limitations
29:59 Takeaways
AI Coffee Break Merch! 🛍️ https://ai-coffee-break-with-letitia....
Thanks to our Patrons who support us in Tier 2, 3, 4: 🙏
Vignesh Valliappan, Ivan Janov, Sunny Dhiana, Andy Ma
▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀
🔥 Optionally, pay us a coffee to help with our Coffee Bean production! ☕
Patreon: / aicoffeebreak
Ko-fi: https://ko-fi.com/aicoffeebreak
Join this channel as a Bean Member to get access to perks:
/ @aicoffeebreak
▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀
🔗 Links:
AICoffeeBreakQuiz: / aicoffeebreak
Twitter / X: / aicoffeebreak
LinkedIn: / letitia-parcalabescu
Threads: https://www.threads.net/@ai.coffee.break
Bluesky: https://bsky.app/profile/aicoffeebrea...
Reddit: / aicoffeebreak
YouTube: / aicoffeebreak
Substack: https://aicoffeebreakwl.substack.com/
Web: https://explanationmark.de/letitia
https://aicoffeebreak.com
#AICoffeeBreak #MsCoffeeBean #MachineLearning #AI #research
Video editing: Nils Trost
Music 🎵 : Space Navigator – Sarah, The Illstrumentarist