Перейти к содержимому

Xiaomi and DeepSeek Just Solved the Biggest Problem in AI

Blake AI Search

0:00 / 0:00

Xiaomi and DeepSeek Just Solved the Biggest Problem in AI

44 просмотра · 9 дн. назад
Blake AI Search
10 подписчиков
44 просмотра · 9 дн. назад
Every step, an AI agent sends its whole history back to the model, and the memory that makes rereading cheap is huge. Twelve days apart, DeepSeek and Xiaomi shrank it with the same two tricks: stop reading halfway, and borrow the list. I walk through Xiaomi's HySparse2 paper and DeepSeek V4.1 Flash, what the numbers really say, who copied whom, and what you can use today, through the API or on your own GPU. Chapters 0:00 The bill 1:00 The memory 2:28 Stop halfway 3:29 Borrow the list 4:31 The numbers 5:30 The twin 6:50 Who copied? 7:40 What you can use Numbers checked on 28 Sep 2026. HySparse2 is an 80 billion parameter research model from a paper, not a released MiMo model. Its 1 million token numbers are the paper's calculations, and "5x less compute" counts operations, not measured time. Prices are DeepSeek's peak rates per 1 million tokens. DeepSeek's bytes per token and Xiaomi's cache size are measured differently, so they are not a head-to-head race. The Hugging Face card, the llama.cpp pull request and the Reddit post are recreated on screen from their pages. The terminal logs are illustrations. Sources Xiaomi · HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing (arXiv 2609.26368): https://arxiv.org/abs/2609.26368 Xiaomi · HySparse (arXiv 2602.03560): https://arxiv.org/abs/2602.03560 Microsoft Research & Tsinghua · YOCO: You Only Cache Once (arXiv 2405.05254): https://arxiv.org/abs/2405.05254 DeepSeek-V4.1-Flash on Hugging Face: https://huggingface.co/deepseek-ai/De... DeepSeek-V4.1-Flash tech report: https://huggingface.co/deepseek-ai/De... DeepSeek API news, Sep 10, 2026: https://api-docs.deepseek.com/news/ne... DeepSeek API pricing: https://api-docs.deepseek.com/quick_s... DeepSeek KV cache guide: https://api-docs.deepseek.com/guides/... OpenRouter on X · 1T tokens on day one: https://x.com/OpenRouter/status/20985... Fuli Luo on X · HySparse2: https://x.com/_LuoFuli/status/2102766... Artificial Analysis · DeepSeek V4.1 Flash: https://artificialanalysis.ai/models/... r/DeepSeek · DeepSeek Flash v4.1, first impressions:   / 1wcfj35   r/LocalLLaMA · HySparse2 thread:   / 1wo7mr6   llama.cpp pull request #28696 · model : add DeepSeek V4.1: https://github.com/ggml-org/llama.cpp... Donato Capitella · DeepSeek V4.1 Flash on Strix Halo:    • DeepSeek V4.1 Flash on Strix Halo: Single-...