Перейти к содержимому

Memory-Efficient LLM Training (Sentic Labcast by Dikshant Kukreja)

senticnet

0:00 / 0:00

Memory-Efficient LLM Training (Sentic Labcast by Dikshant Kukreja)

39 просмотров · 7 дней назад
senticnet
565 подписчиков
39 просмотров · 7 дней назад
Reverse-mode differentiation stores every weight gradient in memory before the optimizer consumes it, creating a large but unnecessary gradient pool. FORGE eliminates this overhead by applying the optimizer directly to each gradient tile in fp32 registers and writing back only the updated weight and moments. This fused step is exact whenever only the optimizer state update reads the gradient and extends naturally to different parallelism strategies. It is architecture agnostic and works across transformers, state-space models, and MLP mixers. FORGE also composes with quantized states, low-rank projections, and factored moments, reducing peak memory by 16–33% while avoiding bf16 gradient truncation. On Llama-3.1-8B, it reduces peak memory from 62.0 GB to 48.4 GB with matched state precision and to 35.3 GB with int8 moments, while achieving 1.5× speedup. It also enables training a 32B model with Muon on a single H200 where standard Muon does not fit. https://sentic.net/memory-efficient-l...