Перейти к содержимому

Every Confusing Thing About How Large Language Models Work Explained Slowly

Roki Explains

0:00 / 0:00

Every Confusing Thing About How Large Language Models Work Explained Slowly

3 603 просмотра · 2 дня назад
Roki Explains
5,1 тыс. подписчиков
3 603 просмотра · 2 дня назад
How do large language models actually work? In this slow, easy-to-follow journey, we explore how AI learned to work with language—from early machine translation and statistical models to word embeddings, attention, Transformers, and modern systems like GPT. We’ll also look at tokenization, next-token prediction, human feedback, hallucinations, and the fascinating question of whether these models truly “understand” anything at all. A calm exploration of the ideas, discoveries, and breakthroughs that turned simple prediction into today’s powerful AI systems. Resources: Attention Is All You Need — Transformer paper https://arxiv.org/abs/1706.03762 Language Models are Few-Shot Learners — GPT-3 https://arxiv.org/abs/2005.14165 Training Language Models to Follow Instructions with Human Feedback https://arxiv.org/abs/2203.02155 Constitutional AI: Harmlessness from AI Feedback https://arxiv.org/abs/2212.08073