A 35B Model on a Mac Mini — Streaming Its Experts Off the SSD
AI deepdive
0:00 / 0:00
A 35B Model on a Mac Mini — Streaming Its Experts Off the SSD
23 просмотра · 11 дн. назад
AI deepdive
10 подписчиков
23 просмотра · 11 дн. назад
Edge0 runs a 35-billion-parameter mixture-of-experts model on a Mac mini M4 Pro at 14.9–17.7 tokens per second, with peak memory under 3 GB. Not by shrinking the model — by streaming the experts off the SSD.
This video walks through the three mechanisms that make it work:
1. SSD expert offload — peak memory is bounded by the active set, not the parameter count.
2. A trained prerouter — predicts routing one step ahead so disk reads overlap the forward pass (+59% decode throughput).
3. Recover-LoRA — 4-bit adapters distilled from the FP teacher, recovering an average 3.9 points across AIME, HumanEval, GPQA-Diamond, MMLU-Pro, and IFBench.
Then the honest edges: cold-start disk faults, Apple Silicon only for now, KV cache still grows with context, and the prerouter is a heuristic — +59% is an average, not a floor.
The reframe: for a sparse MoE, disk is a memory tier now. "Does it fit" is a question about your SSD.
Links:
• Repo: https://github.com/Edge0-AI/Edge0
• Edge0-35B-A3B-preview: https://huggingface.co/Edge0/Edge0-35B-A3B...
• Edge0-8B-A1B-preview: https://huggingface.co/Edge0/Edge0-8B-A1B-...
Chapters:
00:00 A 35B model on a Mac mini
00:48 Why MoE lets this work
01:30 Mechanism 1 — SSD expert offload
02:39 Mechanism 2 — the prerouter
03:47 Mechanism 3 — Recover-LoRA
04:55 The honest edges
05:55 Disk is a memory tier now
#LLM #MoE #LocalLLM #AppleSilicon #Inference