Перейти к содержимому

Qwen 3.8 Flash Next (Fully Tested & VS GLM-5.3 Flash): You can run it LOCALLY! & IT"S CRAZY!

AICodeKing

0:00 / 0:00

Qwen 3.8 Flash Next (Fully Tested & VS GLM-5.3 Flash): You can run it LOCALLY! & IT"S CRAZY!

15 179 просмотров · 19 часов назад
AICodeKing
132 тыс. подписчиков
15 179 просмотров · 19 часов назад
Visit AISeeKing (my second channel with more cool local ai stuff):    / @aiseeking   In this video, I'll be testing Alibaba’s new Qwen 3.8 Flash Next model, which is an experimental preview of the upcoming Qwen 4 architecture. It’s a 125-billion-parameter MoE model that only activates 6 billion parameters per token, making it extremely cheap and efficient, while still offering strong agentic, math, and long-context performance. -- Key Takeaways: 🚀 Alibaba released Qwen 3.8 Flash Next as an experimental preview of the future Qwen 4 architecture. ⚡ It is a 125B parameter MoE model, but only activates 6B parameters per token for very low compute cost. 💸 The API pricing is extremely cheap at 16 cents per million input tokens and 47 cents per million output tokens. 🧠 The model uses a hybrid architecture with Gated DeltaNet blocks and Qwen Sparse Attention for long-context efficiency. 📚 Around 51B parameters are used for bigram and trigram embeddings, which can sit in regular RAM instead of GPU memory. 📊 On KingBench, Qwen 3.8 Flash Next scored 56 out of 80, while GLM 5.3 Flash scored 63 out of 80. ✅ It performed very well on math and agentic tasks, getting perfect scores in both areas. 🎨 Its main weakness was frontend, visual, and 3D generation, where GLM 5.3 Flash performed better. 💻 The model can be run locally through Ollama, llama.cpp, vLLM, SGLang, and GGUF quants from Unsloth. 👍 Overall, Qwen 3.8 Flash Next is not the best main coding model, but it is a very strong cheap model for agentic workflows, tool use, and long-context tasks.