Before You Run AI Locally, Watch This
Cloud Codes
0:00 / 0:00
Before You Run AI Locally, Watch This
7 075 просмотров · 1 день назад
Cloud Codes
40,7 тыс. подписчиков
7 075 просмотров · 1 день назад
A 5.2GB model on a 6GB GPU sounds like an easy fit until it crashes or quietly spills onto your CPU. In this 30-minute guide, we break down why download size is only part of your memory bill, how the KV cache inflates with context, and how to pick the right local setup for your hardware.
🔔 Subscribe: / @cloud-codes
💙 Become a Member: / @cloud-codes
🐦 Twitter/X:
https://x.com/cloud_codes
💬 Discord:
/ discord
In this 30-minute deep dive, We breaks down how LLMs actually run on consumer hardware.
We covered:
• Model weights vs KV cache and working memory
• Dense vs MoE architectures
• GGUF quantization and bits-per-weight
• Ollama context allocation and VRAM usage
• TTFT vs token generation speed
• Ollama, LM Studio, llama.cpp, MLX, Open WebUI, vLLM, SGLang & FreeToken
• 8 performance levers: layer offloading, context sizing, quantization, Flash Attention, KV-cache quantization, prompt caching, speculative decoding & MTP
• RAM offloading, MoE streaming, multi-GPU splits and SSD streaming
• Real builds for 6GB, 24GB and 64GB systems
Build, solve, deploy.
⚙️ Quick Commands:
• Check GPU/CPU model split: ollama ps
• Estimate LM Studio load: lms load --estimate-only model_name --context-length tokens
• Scan hardware with llmfit: brew install AlexsJones/tap/llmfit && llmfit
🔗 Resources Mentioned:
• Ollama GPU & Memory FAQ:
https://docs.ollama.com/faq
• Ollama Context Length Tiers:
https://docs.ollama.com/context-length
• LM Studio CLI Load & Estimation:
https://lmstudio.ai/docs/cli/local-mo...
• LM Studio Speculative Decoding:
https://lmstudio.ai/docs/app/advanced...
• llama.cpp Server & Offloading Knobs:
https://github.com/ggml-org/llama.cpp...
• llama.cpp Multi-GPU Modes (Layer vs. Tensor):
https://github.com/ggml-org/llama.cpp...
• Apple MLX & MLX-LM:
https://github.com/ml-explore/mlx-lm
• Hugging Face GGUF Specification:
https://huggingface.co/docs/hub/gguf
• Hugging Face Llama 3.1 Memory Breakdown:
https://huggingface.co/blog/llama31
• llmfit System Estimator:
https://github.com/AlexsJones/llmfit
• vLLM GPU Installation & MTP Specs:
https://docs.vllm.ai/en/latest/gettin...
• FreeToken Paper (MoE Cross-Tier Inference):
https://arxiv.org/abs/2608.16157
• DeepSeek V4.1 Flash SSD Streaming (antirez):
https://huggingface.co/antirez/deepse...
⏱️ Chapters:
0:00 - The 5.2GB Trap: Why Models Spill & Crawl
1:24 - Reading Model Specs: Parameters, Dense vs. MoE
3:02 - Quantization & GGUF: Bits Per Weight Explained
4:46 - The Three Memory Types: VRAM, RAM & Unified Memory
6:02 - The True Memory Bill: Weights, KV Cache & Overhead
7:11 - Ollama's Hidden Context Trap & VRAM Bands
8:47 - Testing Capacity: 6GB Laptop, 24GB Desktop, 64GB Mac
10:44 - The Two Kinds of Slow: TTFT vs. Generation (tok/s)
12:18 - Software Showdown: Ollama, LM Studio, llama.cpp & MLX
15:02 - Server & Research Stack: vLLM, SGLang, FreeToken & Open WebUI
17:13 - Hardware Realities: Windows/Linux GPUs vs. Apple Silicon
19:15 - Multi-GPU Scaling: Layer Split vs. Tensor Split
19:57 - 8 Levers to Make Local AI Run Faster
22:28 - Speculative Decoding & Multi-Token Prediction (MTP)
23:47 - Beyond VRAM: RAM Offloading & MoE Expert Streaming
25:34 - Extreme SSD Streaming: 755B Models on Local Storage
27:33 - The Right Way to Choose: llmfit & Pre-Load Estimators
28:20 - Final Setups: 6GB Laptop, 24GB Desktop & 64GB Mac
30:20 - The 8-Step Pre-Flight Checklist
31:08 - Summary: What We Learned Today
32:03 - The Horizon Question & Closing Thoughts
#localai #ollama #lmstudio #llamacpp #gpu #vram #cloudcodes
User Queries:
why does my local ai model run so slow
how much vram do i need to run llama 3
ollama model fits in vram but uses cpu
ollama ps processor 50 percent cpu 50 percent gpu
difference between dense model and mixture of experts
how to calculate kv cache memory size
what does q4_k_m mean in gguf
lm studio vs ollama vs llama cpp which is best
how to run large llm on 16gb ram
can i run deepseek locally on macbook
ollama change default context length
speculative decoding lm studio setup
what is llmfit and how to use it
multi gpu layer split vs tensor split llama cpp
deepseek v4.1 flash ssd streaming antirez