Перейти к содержимому

Literally Everything About How Large Language Models Work Explained Slowly (For Sleep)

Cosmo Explains

0:00 / 0:00

Literally Everything About How Large Language Models Work Explained Slowly (For Sleep)

7 992 просмотра · 1 день назад
Cosmo Explains
25,8 тыс. подписчиков
7 992 просмотра · 1 день назад
Now streaming on Spotify https://open.spotify.com/show/033FVSz... Large language models can write, translate, answer questions, and generate code—but how do they actually work? This video explores the story and mechanics behind modern LLMs, from early neural networks, backpropagation, word embeddings, and recurrent networks to attention, Transformers, GPT, BERT, scaling laws, tokenization, and RLHF. We also look at how these models generate text, why they hallucinate, and the challenges surrounding reasoning, alignment, safety, and human oversight. Resources: Attention Is All You Need — Vaswani et al. https://arxiv.org/abs/1706.03762 BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding — Devlin et al. https://arxiv.org/abs/1810.04805 Language Models are Few-Shot Learners — Brown et al. https://arxiv.org/abs/2005.14165 Training Language Models to Follow Instructions with Human Feedback — Ouyang et al. https://arxiv.org/abs/2203.02155