Перейти к содержимому

Flash in the Package Is a Read Budget, Not a Capacity Upgrade

Kemmu Draws Tech

0:00 / 0:00

Flash in the Package Is a Read Budget, Not a Capacity Upgrade

5 просмотров · 7 дней назад
Kemmu Draws Tech
9 подписчиков
5 просмотров · 7 дней назад
HBFSim argues that HBF capacity and data-placement decisions have to be settled before silicon exists, because no storage replay, GPU simulator, or cycle-accurate model can run the real kernels against a flash tier. This episode takes the pitch at face value and does the decode-step arithmetic: how many bytes a mixture-of-experts model pulls per token, how many of those a flash tier can absorb before decode latency moves, and how often placement is allowed to change before NAND write endurance becomes the binding constraint. Chapters 0:00 The pitch 0:39 The budget 1:13 What a token costs 1:59 19.2 milliseconds 2:22 The capacity axis does not reach 3:09 Three pipes 3:45 Heat spends your writes 4:11 The rewrite budget 4:48 Three tools that cannot answer 5:21 Rewriting PTX 6:06 55 to 1 6:44 Sweep economics 7:26 The whole thing on one line 7:47 Allowances, not capacity Sources HBFSim: Fast and Faithful Simulation of High-Bandwidth Flash Under Real GPU Execution — https://arxiv.org/abs/2609.09800v1 LLM in a flash:Efficient Large Language Model Inference with Limited Memory — https://arxiv.org/abs/2312.11514 DeepSeek-V3 Technical Report — https://arxiv.org/abs/2412.19437 HBM3E | Micron Technology Inc. — https://www.micron.com/products/memor... GB200 NVL72 | NVIDIA — https://www.nvidia.com/en-us/data-cen... Drawn with Inkstack.