How does a Large Language Model actually work?
Lag Odhuu
0:00 / 0:00
How does a Large Language Model actually work?
707 просмотров · 2 недели назад
Lag Odhuu
371 подписчик
707 просмотров · 2 недели назад
Ee episode lo memu 3Blue1Brown yokka "Large Language Models explained briefly" lesson ni line by line chaduvutunnam — multi-layer perceptron nunchi modalu petti, oka LLM asalu next word ni ela predict chestundo akkada daaka.
We cover why a model picks a less likely word on purpose, why the same question gets different answers from Claude and GPT, how unlabelled internet text turns out to be labelled after all, why this needs GPUs and how that made Nvidia, RLHF and the whole industry it created, embeddings, and a first look at attention. Transformers deep dive is the next episode.
ABOUT THE SOURCE MATERIAL
This episode follows and discusses Grant Sanderson's (3Blue1Brown) lesson
"Large Language Models explained briefly" — https://www.3blue1brown.com/lessons/m...
We read from it on screen and talk through it in Telugu, adding our own
explanations, disagreements and worked examples. The use is for criticism,
commentary and education. We do not own the original work, and all rights in
it remain with 3Blue1Brown. Please watch the original — it is excellent, and
this episode is a companion to it, not a replacement.
CHAPTERS
00:00 Where we left off — the multi-layer perceptron
01:03 An LLM is a next-word predictor
02:50 Probability, not certainty — why it picks a less likely word
04:16 Why Claude and GPT answer differently
06:08 Training as tuning billions of dials
08:46 "The next word is the label" — how unlabelled text trains a model
10:26 Backpropagation, and why that one sentence now makes sense
12:02 Running GPT-2 on a laptop vs GPT-3
14:07 The compute wall, and how Nvidia became the AI company
16:49 Why a transformer trains in parallel and an RNN cannot
19:01 Pre-training is only half of it
20:37 RLHF, and the industry it created
23:11 Mercor, and paying domain experts to grade models
24:05 Reinforcement learning — penalties, points and gamification
26:26 2017 — Google's transformer reads the whole text at once
27:21 Embeddings — turning a word into a vector of meaning
30:29 Cosine distance, and why "bank" lives in two places
31:00 Attention — letting the words talk to each other
35:12 Feed forward, and stacking attention on attention
38:17 Emergence — why nobody can say why a model chose what it chose
39:18 Interpretability, and Anthropic tracing the thoughts of an LLM
41:23 Next episode — a deep dive into transformers
EVERYTHING WE REFERENCED
3Blue1Brown, Large Language Models explained briefly — https://www.3blue1brown.com/lessons/m...
3Blue1Brown, Neural networks series — https://www.3blue1brown.com/?topic=ne...
3Blue1Brown, Analyzing our neural network — https://www.3blue1brown.com/lessons/n...
Attention Is All You Need (Vaswani et al., 2017) — https://arxiv.org/abs/1706.03762
Anthropic, Tracing the thoughts of a large language model — https://www.anthropic.com/research/tr...
Anthropic, Circuit Tracing: Revealing Computational Graphs in Language Models — https://transformer-circuits.pub/2025...
Anthropic, On the Biology of a Large Language Model — https://transformer-circuits.pub/2025...
LessWrong discussion of the tracing-thoughts work — https://www.lesswrong.com/posts/zsr4r...
Mercor — https://www.mercor.com
Mercor for domain experts — https://www.mercor.com/experts/
Scale AI — https://scale.com
Perplexity — https://www.perplexity.ai
Word2Vec (Mikolov et al., 2013) — https://arxiv.org/abs/1301.3781
GloVe, Stanford NLP — https://nlp.stanford.edu/projects/glove/
NVIDIA DGX Spark — https://www.nvidia.com/en-us/products...
Subtext — https://github.com/ninjahawk/Subtext
🌟 Building the new age Telugu tech ecosystem 🌟
Follow us: https://linktr.ee/lagodhuu
Instagram: / lagodhuu
#Telugu #TeluguTech #LLM #Transformers #NeuralNetworks #AttentionMechanism #MachineLearning #3Blue1Brown #LagOdhuu