AMD's New Mini PC Holds 192GB Of AI Memory
The Stack
0:00 / 0:00
AMD's New Mini PC Holds 192GB Of AI Memory
18 958 просмотров · 1 день назад
The Stack
11,8 тыс. подписчиков
18 958 просмотров · 1 день назад
Ryzen AI Max+ vs DGX Spark: $700 cheaper, 192GB unified memory, but AMD loses on prompt processing speed, here's which local AI desktop actually wins.
AMD's Ryzen AI Halo workstation costs $3,999 against NVIDIA's $4,699 DGX Spark, yet the real difference isn't price, it's how unified memory actually changes the game for running large language models locally. Both machines cap at 128GB of memory today, but AMD's architecture shares that pool between CPU and integrated GPU, letting you load 70B+ parameter models that would never fit in a discrete graphics card's VRAM. The newer Ryzen AI Max+ PRO 495 chips shown at IFA 2026 push that ceiling to 192GB, though neither the Acemagic F9A-Pro495 nor Minisforum MS-S1 MAX-P495 have launched pricing or dates yet.
Here's the catch: on Linux, the GPU only sees about 108GB of that 128GB pool, and memory bandwidth sits at 256GB/s on both AMD and NVIDIA, nearly identical to the DGX Spark's 273GB/s. Writing text? Both machines generate a 235B-parameter model at roughly 11 tokens per second. But reading a prompt (processing your input) is where AMD falls apart: 168 tokens/sec on the Ryzen versus 2,100 on the DGX Spark. That gap explodes exponentially with longer context windows, making the headline 192GB feature essentially more room to be slow in. On token generation alone, you save $700 and match a $9,398 dual-Spark setup. But if your workflow demands heavy prompt processing, codebases, long PDFs, transcripts, that savings evaporates fast.
This breakdown is for builders choosing between local AI on unified-memory mini-PCs (Minisforum, Acemagic boxes, AMD Ryzen AI Halo) versus NVIDIA's workstations, and anyone deciding whether a $3,999 or $8,000 machine actually fits their inference needs.
Chapters:
0:00 $700 cheaper, but what's the real difference?
0:46 How shared memory actually changes everything
2:12 The specs nobody saw coming
3:50 Why 128GB becomes 108GB on Linux
4:54 Two hundred billion parameters at home
6:07 TOPS numbers that don't matter
7:11 Two liters, 126 TOPS, one hand
8:20 Memory bandwidth: the real bottleneck
9:19 Where NVIDIA's extra cost shows up
10:11 When two Sparks beat one box
11:38 The one port worth $700
12:50 Reading texts: AMD's invisible penalty
14:11 Context windows expose the weakness
15:07 More room to be slow in
16:18 The 192GB machine's shocking price
17:57 Better for chat, worse for coding
19:03 Your workflow decides everything
Tools & resources mentioned:
Ryzen AI Halo (AMD Developer Workstation): https://www.newline.co/@Dipen/use-loc...
Minisforum MS-S1 Max: https://www.amazon.com/MINISFORUM-AMD...
Acemagic F9A Mini-Workstation: https://acemagic.eu/blogs/einkaufsfue...
NVIDIA DGX Spark: https://forums.developer.nvidia.com/t...
Open WebUI: https://openwebui.com
Ollama: https://ollama.ai
llama.cpp: https://github.com/ggerganov/llama.cpp
About The Stack
The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs.
We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship.
Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai...
#local ai #unified memory #llm inference #mini pc #dgx spark alternative