Literally Everything About How Large Language Models Work Explained Slowly (For Sleep)
Cosmo Explains
0:00 / 0:00
Literally Everything About How Large Language Models Work Explained Slowly (For Sleep)
7 992 просмотра · 1 день назад
Cosmo Explains
25,8 тыс. подписчиков
7 992 просмотра · 1 день назад
Now streaming on Spotify
https://open.spotify.com/show/033FVSz...
Large language models can write, translate, answer questions, and generate code—but how do they actually work?
This video explores the story and mechanics behind modern LLMs, from early neural networks, backpropagation, word embeddings, and recurrent networks to attention, Transformers, GPT, BERT, scaling laws, tokenization, and RLHF. We also look at how these models generate text, why they hallucinate, and the challenges surrounding reasoning, alignment, safety, and human oversight.
Resources:
Attention Is All You Need — Vaswani et al.
https://arxiv.org/abs/1706.03762
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding — Devlin et al.
https://arxiv.org/abs/1810.04805
Language Models are Few-Shot Learners — Brown et al.
https://arxiv.org/abs/2005.14165
Training Language Models to Follow Instructions with Human Feedback — Ouyang et al.
https://arxiv.org/abs/2203.02155