Qwen3.8-flash-next Killed the VRAM Ceiling? Not Quite.
The Local Ceiling
0:00 / 0:00
Qwen3.8-flash-next Killed the VRAM Ceiling? Not Quite.
1 461 просмотр · 4 недели назад
The Local Ceiling
153 подписчика
1 461 просмотр · 4 недели назад
The 113 GB download of Qwen3.8-Flash-Next decodes FASTER than the 94 GB
one — same weights, at every context depth we tested. That's not the
interesting part. The interesting part is why, because the why decides
which quant you should be downloading tonight. So how should you pick your quant?
The internet spent a month saying this model's native sparse attention
killed the VRAM ceiling. We ran it locally — CPU experts, one datacenter
GPU — and measured where the ceiling actually is. It isn't your VRAM. On
this machine it isn't even the memory bandwidth, and we'll show you the
one measured number that rules the usual suspect out.
Then we point a survey tool we're building at this model in its release
week — a model it had never seen — and let it find the ceiling on its
own.
All runs on our own hardware (Dell R730, Tesla V100 32GB + P40 + P4),
stock llama.cpp, stock published quants, 2026-08-29 → 08-31. Full bench
records, every command, and the chart data:
https://github.com/firstpartyworks/ke...
Watch next:
The VRAM Illusion (fit vs spill — the prequel): • The VRAM Illusion: Why gpt-oss-120b's Cost...
Stop Blindly Quantizing Your KV Cache: • Stop Blindly Quantizing Your KV Cache (We ...
#Qwen3 #LocalLLM #llamacpp
0:00 113 GB beats 94 GB
0:57 Sign-on
1:04 The claim
2:13 What it actually runs like
3:19 The myth-check verdict
4:20 Not the pipe
5:15 It's the kernels
6:16 The proof: the bigger file wins
7:39 Qwen already knew
8:20 Surveying a brand-new model
9:32 Taught once, runs solo
11:00 Where your ceiling is