Andrej Karpathy at Stanford: The Foundation that Became Anthropic
DarrowStation
0:00 / 0:00
Andrej Karpathy at Stanford: The Foundation that Became Anthropic
16 454 просмотра · 2 недели назад
DarrowStation
350 подписчиков
16 454 просмотра · 2 недели назад
Why Frontier Reasoning Models Scale Test-Time Compute: The Complete Masterclass
An expert masterclass on how modern frontier AI models achieve step-by-step reasoning and complex problem-solving. In this focused 1-hour deep dive, we explore why post-training is replacing pure pre-training scale, how inference-time search works with Process Reward Models (PRMs), and how to apply these architectural principles to your own systems.
⏳ Chapters:
00:00 - Masterclass Introduction: The Test-Time Compute Paradigm
05:30 - Pre-Training Limits & The Shift to System 2 Inference
14:15 - Reward Modeling: Outcome (ORM) vs. Process-Based (PRM)
25:40 - Reinforcement Learning Alignment: PPO, DPO & GRPO
36:10 - Search Tree Architectures & MCTS at Inference Time
48:25 - Scaling Limits, Context Window Overhead & Evaluation
📝 About this Video:
In this masterclass, we dissect the paradigm shift from traditional training-time scaling laws to test-time compute and inference-time reasoning architectures. By studying the latest post-training patterns—including Process-Supervised Reward Models (PRMs), Group Relative Policy Optimization (GRPO), and search-guided token generation—engineers and AI researchers can understand how to build, evaluate, and deploy reasoning-grade AI systems.
🔗 Resources Mentioned:
Subscribe on YouTube: @darrowstation
Don't forget to Subscribe for more deep dives into AI engineering!#ai #machinelearning #LLMs #ReinforcementLearning #DeepLearning #softwareengineering #posttraining