Every Confusing Thing About How Large Language Models Work Explained Slowly
Roki Explains
0:00 / 0:00
Every Confusing Thing About How Large Language Models Work Explained Slowly
3 603 просмотра · 2 дня назад
Roki Explains
5,1 тыс. подписчиков
3 603 просмотра · 2 дня назад
How do large language models actually work?
In this slow, easy-to-follow journey, we explore how AI learned to work with language—from early machine translation and statistical models to word embeddings, attention, Transformers, and modern systems like GPT. We’ll also look at tokenization, next-token prediction, human feedback, hallucinations, and the fascinating question of whether these models truly “understand” anything at all.
A calm exploration of the ideas, discoveries, and breakthroughs that turned simple prediction into today’s powerful AI systems.
Resources:
Attention Is All You Need — Transformer paper
https://arxiv.org/abs/1706.03762
Language Models are Few-Shot Learners — GPT-3
https://arxiv.org/abs/2005.14165
Training Language Models to Follow Instructions with Human Feedback
https://arxiv.org/abs/2203.02155
Constitutional AI: Harmlessness from AI Feedback
https://arxiv.org/abs/2212.08073