Перейти к содержимому

180B Model. 12GB GPU. What Speed Do Real PCs Get?

Blake AI Search

0:00 / 0:00

180B Model. 12GB GPU. What Speed Do Real PCs Get?

71 просмотр · 9 дн. назад
Blake AI Search
10 подписчиков
71 просмотр · 9 дн. назад
A 180 billion parameter model on a 12 GB graphics card, at 46 tokens a second. That was one developer's own PC. I went through what other people got with the same free app, Strata, and the same model, Qwen3.8-Flash-Next: from 17 tokens a second on a gaming laptop to 118 on an RTX 5090. Here's how a model this big fits on a small card, what actually sets your speed, the five catches, and what I'd do on a desktop, a laptop, or a PC with 32 GB of RAM. Chapters 0:00 The question 0:57 How it fits 2:43 The developer's desk 3:41 Real PCs 6:29 The catches 8:14 Is it worth it? 9:30 My call Numbers checked on 29 Sep 2026. The speeds come from different cards, model versions and conversation lengths, reported by their owners: they show the spread, not which part caused it. This is not a lab test, and I did not measure them myself. The developer's numbers are from the Strata docs. A token is about 3/4 of a word (Strata README), so the "speed demo" boxes type at the measured rate; the text in them is only an example. LiveCodeBench scores (87.43, 86.29 and 81.14) are from the ISTA-DASLab model card. SWE-bench Pro is Qwen's own chart. The Intelligence Index is from Artificial Analysis. Cloud prices for Qwen3.8 Flash as listed in late September 2026. The BIOS, Windows and settings screens are illustrations. Quotes from GitHub and Reddit are shown as their authors wrote them. Sources Qwen blog · Qwen3.8-Flash-Next: https://qwen.ai/blog?id=qwen3.8-flash... Qwen3.8-Flash-Next model card (Hugging Face): https://huggingface.co/Qwen/Qwen3.8-F... Qwen3.8-Flash-Next license: https://huggingface.co/Qwen/Qwen3.8-F... Strata (GitHub, Niko1221): https://github.com/Niko1221/Strata Strata docs · DETAILS.md: https://github.com/Niko1221/Strata/bl... Strata license history: https://github.com/Niko1221/Strata/co... Strata demo · Pagoda.mp4: https://github.com/Niko1221/Strata/re... Strata issue #74 · laptop RTX 3080, RTX 5070 with DDR4, XMP: https://github.com/Niko1221/Strata/is... Strata issue #71 · tuned llama.cpp fork, RTX 5070 Ti: https://github.com/Niko1221/Strata/is... Strata issue #60 · out of memory on Windows: https://github.com/Niko1221/Strata/is... jun76 · Qwen3.8-Flash-Next measurements on an RTX 5090: https://github.com/jun76/qwen3.8-flas... r/LocalLLaMA · the developer's thread:   / qwen38flashnext_on_12gb_vram_65_tokens_per...   r/LocalLLaMA · mini PC with an RTX 3060 over OCuLink:   / pc5ink3   r/LocalLLaMA · RTX 5080 results (Sid3effect):   / pc66ie0   r/LocalLLaMA · an early tester on sloppy prompts and the install (blackal1ce):   / pby6u57   r/LocalLLaMA · a question whether the answers match a normal engine:   / pbt9p54   ISTA-DASLab · Qwen3.8-Flash-Next GSQ-RCO GGUF: https://huggingface.co/ISTA-DASLab/Qw... Artificial Analysis · Qwen3.8-Flash-Next: https://artificialanalysis.ai/models/... Artificial Analysis · model leaderboard: https://artificialanalysis.ai/leaderb...