Перейти к содержимому

Before You Run AI Locally, Watch This

Cloud Codes

0:00 / 0:00

Before You Run AI Locally, Watch This

7 075 просмотров · 1 день назад
Cloud Codes
40,7 тыс. подписчиков
7 075 просмотров · 1 день назад
A 5.2GB model on a 6GB GPU sounds like an easy fit until it crashes or quietly spills onto your CPU. In this 30-minute guide, we break down why download size is only part of your memory bill, how the KV cache inflates with context, and how to pick the right local setup for your hardware. 🔔 Subscribe:    / @cloud-codes   💙 Become a Member:    / @cloud-codes   🐦 Twitter/X: https://x.com/cloud_codes 💬 Discord:   / discord   In this 30-minute deep dive, We breaks down how LLMs actually run on consumer hardware. We covered: • Model weights vs KV cache and working memory • Dense vs MoE architectures • GGUF quantization and bits-per-weight • Ollama context allocation and VRAM usage • TTFT vs token generation speed • Ollama, LM Studio, llama.cpp, MLX, Open WebUI, vLLM, SGLang & FreeToken • 8 performance levers: layer offloading, context sizing, quantization, Flash Attention, KV-cache quantization, prompt caching, speculative decoding & MTP • RAM offloading, MoE streaming, multi-GPU splits and SSD streaming • Real builds for 6GB, 24GB and 64GB systems Build, solve, deploy. ⚙️ Quick Commands: • Check GPU/CPU model split: ollama ps • Estimate LM Studio load: lms load --estimate-only model_name --context-length tokens • Scan hardware with llmfit: brew install AlexsJones/tap/llmfit && llmfit 🔗 Resources Mentioned: • Ollama GPU & Memory FAQ: https://docs.ollama.com/faq • Ollama Context Length Tiers: https://docs.ollama.com/context-length • LM Studio CLI Load & Estimation: https://lmstudio.ai/docs/cli/local-mo... • LM Studio Speculative Decoding: https://lmstudio.ai/docs/app/advanced... • llama.cpp Server & Offloading Knobs: https://github.com/ggml-org/llama.cpp... • llama.cpp Multi-GPU Modes (Layer vs. Tensor): https://github.com/ggml-org/llama.cpp... • Apple MLX & MLX-LM: https://github.com/ml-explore/mlx-lm • Hugging Face GGUF Specification: https://huggingface.co/docs/hub/gguf • Hugging Face Llama 3.1 Memory Breakdown: https://huggingface.co/blog/llama31 • llmfit System Estimator: https://github.com/AlexsJones/llmfit • vLLM GPU Installation & MTP Specs: https://docs.vllm.ai/en/latest/gettin... • FreeToken Paper (MoE Cross-Tier Inference): https://arxiv.org/abs/2608.16157 • DeepSeek V4.1 Flash SSD Streaming (antirez): https://huggingface.co/antirez/deepse... ⏱️ Chapters: 0:00 - The 5.2GB Trap: Why Models Spill & Crawl 1:24 - Reading Model Specs: Parameters, Dense vs. MoE 3:02 - Quantization & GGUF: Bits Per Weight Explained 4:46 - The Three Memory Types: VRAM, RAM & Unified Memory 6:02 - The True Memory Bill: Weights, KV Cache & Overhead 7:11 - Ollama's Hidden Context Trap & VRAM Bands 8:47 - Testing Capacity: 6GB Laptop, 24GB Desktop, 64GB Mac 10:44 - The Two Kinds of Slow: TTFT vs. Generation (tok/s) 12:18 - Software Showdown: Ollama, LM Studio, llama.cpp & MLX 15:02 - Server & Research Stack: vLLM, SGLang, FreeToken & Open WebUI 17:13 - Hardware Realities: Windows/Linux GPUs vs. Apple Silicon 19:15 - Multi-GPU Scaling: Layer Split vs. Tensor Split 19:57 - 8 Levers to Make Local AI Run Faster 22:28 - Speculative Decoding & Multi-Token Prediction (MTP) 23:47 - Beyond VRAM: RAM Offloading & MoE Expert Streaming 25:34 - Extreme SSD Streaming: 755B Models on Local Storage 27:33 - The Right Way to Choose: llmfit & Pre-Load Estimators 28:20 - Final Setups: 6GB Laptop, 24GB Desktop & 64GB Mac 30:20 - The 8-Step Pre-Flight Checklist 31:08 - Summary: What We Learned Today 32:03 - The Horizon Question & Closing Thoughts #localai #ollama #lmstudio #llamacpp #gpu #vram #cloudcodes User Queries: why does my local ai model run so slow how much vram do i need to run llama 3 ollama model fits in vram but uses cpu ollama ps processor 50 percent cpu 50 percent gpu difference between dense model and mixture of experts how to calculate kv cache memory size what does q4_k_m mean in gguf lm studio vs ollama vs llama cpp which is best how to run large llm on 16gb ram can i run deepseek locally on macbook ollama change default context length speculative decoding lm studio setup what is llmfit and how to use it multi gpu layer split vs tensor split llama cpp deepseek v4.1 flash ssd streaming antirez