Alibaba releases Qwen3.8-Flash-Next, previewing Qwen4 architecture
Qwen3.8-Flash-Next is a 125B multimodal MoE with 51B N-gram embeddings, activating only 6B parameters per token. It beats Claude Opus 4.6 Max on 8 of 9 comparable benchmarks. The model uses Gated DeltaNet and Qwen Sparse Attention for efficient long-context inference.
How this story unfolded
2 days · 7 reports · 12 community posts · 19 of 21 shown
- Aug 25
- Aug 26
Alibaba’s Qwen to open-source Qwen3.8-Flash-Next, previewing Qwen4 architecturetechnode.com
unsloth/Qwen3.8-Flash-Next-GGUFhuggingface.co
Qwen/Qwen3.8-Flash-Next · Hugging Facehuggingface.co
Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecturemarktechpost.com
Qwen/Qwen3.8-Flash-Next-FP8huggingface.co
Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Codingdeveloper.nvidia.com
Qwen by email
Get an email when Qwen has news
No news that day, no email.
More stories today
- Stable Diffusion user tests H3 model with Cheers-style script
- Reddit users share impressive image-to-video AI demos
- Reddit reminds users they can legally seed AI models via torrenting
- MiniMax H3 reverse-engineers paintings into basic forms
- OpenAI DevDay Exchange Seoul applications close Sept 4