Scaling Laws Explained: Why "Just Make It Bigger" Stopped Working
Jimmy's Tech Deep Dive
0:00 / 0:00
Scaling Laws Explained: Why "Just Make It Bigger" Stopped Working
14 просмотров · 6 дней назад
Jimmy's Tech Deep Dive
9 подписчиков
14 просмотров · 6 дней назад
For six years, one idea drove almost all progress in AI: just make it bigger.
Here is the maths behind it, and why the industry is now changing strategy.
A full walkthrough of neural scaling laws, built from nothing. We start with what a power law actually is and why log-log straight lines let labs commit billions to a single training run. Then the two papers that set the rules: Kaplan 2020, which pushed everyone toward huge undertrained models, and DeepMind's Chinchilla in 2022, which showed size and data must grow together at roughly twenty tokens per parameter. From there we look at where the laws break down - exploding costs, the data wall, worn-out benchmarks, architectural limits - and at the new lever labs are pulling instead: test-time compute, where the model thinks longer rather than getting larger.
• N, D, C and loss: the four numbers every scaling law is built from
• Power laws, log-log straight lines, and why diminishing returns were baked in from day one
• Kaplan 2020 vs Chinchilla: how a single exponent misdirected two years of training compute
• IsoFLOP valleys, the 20-tokens-per-parameter rule, and why labs now overtrain on purpose
• The data wall, repeated passes, and why bigger models suffer more from repetition
• Test-time compute, verifiers and search - capability you buy per question
Chapters
0:00 Just make it bigger
1:09 Parameters, data, compute, loss
2:29 What a power law is
4:28 Kaplan 2020
7:01 The undertrained era
7:28 DeepMind checks the maths
8:57 Twenty tokens per parameter
11:37 Why labs overtrain
12:39 Cost and the data wall
15:34 Benchmarks and architecture
16:45 What scaling never promised
17:41 Test-time compute
20:18 Capability per question
21:36 Efficiency gains and loss vs capability
22:58 Four eras, and the fifth
24:30 Why ideas matter again
Sources
Scaling Laws for Neural Language Models - https://arxiv.org/pdf/2001.08361
Training Compute-Optimal Large Language Models - https://proceedings.neurips.cc/paper/...
Chinchilla scaling: A replication attempt (Epoch AI) - https://epoch.ai/publications/chinchi...
Data-Constrained Scaling: Training LLMs Beyond the Data Wall - https://mbrenndoerfer.com/writing/dat...
Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling - https://www.alphaxiv.org/abs/2605.25143
What Is Test-Time Compute? - https://tdwi.org/blogs/ai-101/2026/05...
#ScalingLaws #Chinchilla #TestTimeCompute #LLM #AI