Jing Huang | Why Larger Models Learn More
DatologyAI
0:00 / 0:00
Jing Huang | Why Larger Models Learn More
98 просмотров · 1 месяц назад
DatologyAI
102 подписчика
98 просмотров · 1 месяц назад
Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention | Jing Huang | Summer of Data 2026
Why do bigger models learn rare, complex skills that smaller models miss entirely — even with infinite data? Jing Huang (Stanford) presents research showing smaller models funnel their limited resources toward frequent, simple tasks, starving out rare ones, while larger models sidestep this through reduced gradient interference and better retention of partial progress between sparse task occurrences — revealing that scaling isn't just about raw capacity, but about a fundamentally different competition for resources during training.