Qwen 3.8 Flash Next (Fully Tested & VS GLM-5.3 Flash): You can run it LOCALLY! & IT"S CRAZY!
AICodeKing
0:00 / 0:00
Qwen 3.8 Flash Next (Fully Tested & VS GLM-5.3 Flash): You can run it LOCALLY! & IT"S CRAZY!
15 179 просмотров · 19 часов назад
AICodeKing
132 тыс. подписчиков
15 179 просмотров · 19 часов назад
Visit AISeeKing (my second channel with more cool local ai stuff): / @aiseeking
In this video, I'll be testing Alibaba’s new Qwen 3.8 Flash Next model, which is an experimental preview of the upcoming Qwen 4 architecture. It’s a 125-billion-parameter MoE model that only activates 6 billion parameters per token, making it extremely cheap and efficient, while still offering strong agentic, math, and long-context performance.
--
Key Takeaways:
🚀 Alibaba released Qwen 3.8 Flash Next as an experimental preview of the future Qwen 4 architecture.
⚡ It is a 125B parameter MoE model, but only activates 6B parameters per token for very low compute cost.
💸 The API pricing is extremely cheap at 16 cents per million input tokens and 47 cents per million output tokens.
🧠 The model uses a hybrid architecture with Gated DeltaNet blocks and Qwen Sparse Attention for long-context efficiency.
📚 Around 51B parameters are used for bigram and trigram embeddings, which can sit in regular RAM instead of GPU memory.
📊 On KingBench, Qwen 3.8 Flash Next scored 56 out of 80, while GLM 5.3 Flash scored 63 out of 80.
✅ It performed very well on math and agentic tasks, getting perfect scores in both areas.
🎨 Its main weakness was frontend, visual, and 3D generation, where GLM 5.3 Flash performed better.
💻 The model can be run locally through Ollama, llama.cpp, vLLM, SGLang, and GGUF quants from Unsloth.
👍 Overall, Qwen 3.8 Flash Next is not the best main coding model, but it is a very strong cheap model for agentic workflows, tool use, and long-context tasks.