Flash in the Package Is a Read Budget, Not a Capacity Upgrade
Kemmu Draws Tech
0:00 / 0:00
Flash in the Package Is a Read Budget, Not a Capacity Upgrade
5 просмотров · 7 дней назад
Kemmu Draws Tech
9 подписчиков
5 просмотров · 7 дней назад
HBFSim argues that HBF capacity and data-placement decisions have to be settled before silicon exists, because no storage replay, GPU simulator, or cycle-accurate model can run the real kernels against a flash tier. This episode takes the pitch at face value and does the decode-step arithmetic: how many bytes a mixture-of-experts model pulls per token, how many of those a flash tier can absorb before decode latency moves, and how often placement is allowed to change before NAND write endurance becomes the binding constraint.
Chapters
0:00 The pitch
0:39 The budget
1:13 What a token costs
1:59 19.2 milliseconds
2:22 The capacity axis does not reach
3:09 Three pipes
3:45 Heat spends your writes
4:11 The rewrite budget
4:48 Three tools that cannot answer
5:21 Rewriting PTX
6:06 55 to 1
6:44 Sweep economics
7:26 The whole thing on one line
7:47 Allowances, not capacity
Sources
HBFSim: Fast and Faithful Simulation of High-Bandwidth Flash Under Real GPU Execution — https://arxiv.org/abs/2609.09800v1
LLM in a flash:Efficient Large Language Model Inference with Limited Memory — https://arxiv.org/abs/2312.11514
DeepSeek-V3 Technical Report — https://arxiv.org/abs/2412.19437
HBM3E | Micron Technology Inc. — https://www.micron.com/products/memor...
GB200 NVL72 | NVIDIA — https://www.nvidia.com/en-us/data-cen...
Drawn with Inkstack.