Перейти к содержимому

How does a Large Language Model actually work?

Lag Odhuu

0:00 / 0:00

How does a Large Language Model actually work?

707 просмотров · 2 недели назад
Lag Odhuu
371 подписчик
707 просмотров · 2 недели назад
Ee episode lo memu 3Blue1Brown yokka "Large Language Models explained briefly" lesson ni line by line chaduvutunnam — multi-layer perceptron nunchi modalu petti, oka LLM asalu next word ni ela predict chestundo akkada daaka. We cover why a model picks a less likely word on purpose, why the same question gets different answers from Claude and GPT, how unlabelled internet text turns out to be labelled after all, why this needs GPUs and how that made Nvidia, RLHF and the whole industry it created, embeddings, and a first look at attention. Transformers deep dive is the next episode. ABOUT THE SOURCE MATERIAL This episode follows and discusses Grant Sanderson's (3Blue1Brown) lesson "Large Language Models explained briefly" — https://www.3blue1brown.com/lessons/m... We read from it on screen and talk through it in Telugu, adding our own explanations, disagreements and worked examples. The use is for criticism, commentary and education. We do not own the original work, and all rights in it remain with 3Blue1Brown. Please watch the original — it is excellent, and this episode is a companion to it, not a replacement. CHAPTERS 00:00 Where we left off — the multi-layer perceptron 01:03 An LLM is a next-word predictor 02:50 Probability, not certainty — why it picks a less likely word 04:16 Why Claude and GPT answer differently 06:08 Training as tuning billions of dials 08:46 "The next word is the label" — how unlabelled text trains a model 10:26 Backpropagation, and why that one sentence now makes sense 12:02 Running GPT-2 on a laptop vs GPT-3 14:07 The compute wall, and how Nvidia became the AI company 16:49 Why a transformer trains in parallel and an RNN cannot 19:01 Pre-training is only half of it 20:37 RLHF, and the industry it created 23:11 Mercor, and paying domain experts to grade models 24:05 Reinforcement learning — penalties, points and gamification 26:26 2017 — Google's transformer reads the whole text at once 27:21 Embeddings — turning a word into a vector of meaning 30:29 Cosine distance, and why "bank" lives in two places 31:00 Attention — letting the words talk to each other 35:12 Feed forward, and stacking attention on attention 38:17 Emergence — why nobody can say why a model chose what it chose 39:18 Interpretability, and Anthropic tracing the thoughts of an LLM 41:23 Next episode — a deep dive into transformers EVERYTHING WE REFERENCED 3Blue1Brown, Large Language Models explained briefly — https://www.3blue1brown.com/lessons/m... 3Blue1Brown, Neural networks series — https://www.3blue1brown.com/?topic=ne... 3Blue1Brown, Analyzing our neural network — https://www.3blue1brown.com/lessons/n... Attention Is All You Need (Vaswani et al., 2017) — https://arxiv.org/abs/1706.03762 Anthropic, Tracing the thoughts of a large language model — https://www.anthropic.com/research/tr... Anthropic, Circuit Tracing: Revealing Computational Graphs in Language Models — https://transformer-circuits.pub/2025... Anthropic, On the Biology of a Large Language Model — https://transformer-circuits.pub/2025... LessWrong discussion of the tracing-thoughts work — https://www.lesswrong.com/posts/zsr4r... Mercor — https://www.mercor.com Mercor for domain experts — https://www.mercor.com/experts/ Scale AI — https://scale.com Perplexity — https://www.perplexity.ai Word2Vec (Mikolov et al., 2013) — https://arxiv.org/abs/1301.3781 GloVe, Stanford NLP — https://nlp.stanford.edu/projects/glove/ NVIDIA DGX Spark — https://www.nvidia.com/en-us/products... Subtext — https://github.com/ninjahawk/Subtext 🌟 Building the new age Telugu tech ecosystem 🌟 Follow us: https://linktr.ee/lagodhuu Instagram:   / lagodhuu   #Telugu #TeluguTech #LLM #Transformers #NeuralNetworks #AttentionMechanism #MachineLearning #3Blue1Brown #LagOdhuu