Перейти к содержимому

Q4 vs Q8: What Quantization Really Costs (Qwen 27B)

Ash

0:00 / 0:00

Q4 vs Q8: What Quantization Really Costs (Qwen 27B)

7 048 просмотров · 6 дн. назад
Ash
683 подписчика
7 048 просмотров · 6 дн. назад
Q8 vs Q6 vs Q5 vs Q4 vs Q3 vs Q2: how much quality does LLM quantization really cost? Measured on Qwen 3.8 27B GGUF, plus KV cache quantization and which file to download for 16, 24, 32 and 48GB Macs. Qwen 3.8 27B is 54 GB at 16-bit. 8-bit is 29 GB, 4-bit (Q4_K_M) 17 GB, 3-bit 14 GB, 2-bit 9 GB. In one publisher's test, the 8-bit file picks the same next word as the original 98.9% of the time, 4-bit 95.6%, 3-bit 92.4% and the smallest 2-bit 79.4%: almost flat down to 4-bit, then a cliff below about 14 GB. Errors stack: 0.956 to the power of 100 is about 1%. 2025 studies found 4-bit weights near lossless on reasoning models but lower bit-widths risky, and up to 59% loss for some 4-bit methods on 64K+ token inputs (older models). The hidden memory hog is the KV cache: about 256 KB per token on Qwen 3.8 27B, 2 GB at 8K context and 8 GB at 32K. An 8-bit KV cache halves it with tiny measured damage on Qwen, but Gemma 4 is far more sensitive. Then: the best quant for your RAM. Subscribe for daily breakdowns of local AI with real numbers, and tell me in the comments which quant you run and on what hardware. Chapters: 0:00 Intro 1:02 What quantization trades 2:37 Why rounding hurts 3:42 Measuring the damage 5:00 Why one wrong word matters 6:20 How to read GGUF file names 7:04 Same name, different file 8:06 Bigger model or more bits? 8:46 The hidden memory hog 10:22 Which file for your RAM 11:38 Three rules 12:11 The receipt #LocalAI #LLM #Quantization