Перейти к содержимому

Qwen3.8-flash-next Killed the VRAM Ceiling? Not Quite.

The Local Ceiling

0:00 / 0:00

Qwen3.8-flash-next Killed the VRAM Ceiling? Not Quite.

1 461 просмотр · 4 недели назад
The Local Ceiling
153 подписчика
1 461 просмотр · 4 недели назад
The 113 GB download of Qwen3.8-Flash-Next decodes FASTER than the 94 GB one — same weights, at every context depth we tested. That's not the interesting part. The interesting part is why, because the why decides which quant you should be downloading tonight. So how should you pick your quant? The internet spent a month saying this model's native sparse attention killed the VRAM ceiling. We ran it locally — CPU experts, one datacenter GPU — and measured where the ceiling actually is. It isn't your VRAM. On this machine it isn't even the memory bandwidth, and we'll show you the one measured number that rules the usual suspect out. Then we point a survey tool we're building at this model in its release week — a model it had never seen — and let it find the ceiling on its own. All runs on our own hardware (Dell R730, Tesla V100 32GB + P40 + P4), stock llama.cpp, stock published quants, 2026-08-29 → 08-31. Full bench records, every command, and the chart data: https://github.com/firstpartyworks/ke... Watch next: The VRAM Illusion (fit vs spill — the prequel):    • The VRAM Illusion: Why gpt-oss-120b's Cost...   Stop Blindly Quantizing Your KV Cache:    • Stop Blindly Quantizing Your KV Cache (We ...   #Qwen3 #LocalLLM #llamacpp 0:00 113 GB beats 94 GB 0:57 Sign-on 1:04 The claim 2:13 What it actually runs like 3:19 The myth-check verdict 4:20 Not the pipe 5:15 It's the kernels 6:16 The proof: the bigger file wins 7:39 Qwen already knew 8:20 Surveying a brand-new model 9:32 Taught once, runs solo 11:00 Where your ceiling is