Alibaba releases Qwen3.8-Flash-Next, previewing Qwen4 architecture
Qwen3.8-Flash-Next is a multimodal MoE with 125B parameters plus 51B N-gram embeddings, activating only 6B per token. It beats Claude Opus 4.6 Max on 8 of 9 comparable benchmarks. QwenCloud API pricing: $0.16/1M input and $0.47/1M output tokens.
How this story unfolded
3 days · 11 reports · 26 community posts · 37 of 40 shown
- Aug 25
- Aug 26
Alibaba’s Qwen to open-source Qwen3.8-Flash-Next, previewing Qwen4 architecturetechnode.com
unsloth/Qwen3.8-Flash-Next-GGUFhuggingface.co
Qwen/Qwen3.8-Flash-Next · Hugging Facehuggingface.co
Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecturemarktechpost.com
Qwen/Qwen3.8-Flash-Next-FP8huggingface.co
Qwen 3.8 Flash now available on AI Gatewayvercel.com
Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Codingdeveloper.nvidia.com
- Aug 27
- Aug 28
Qwen by email
Get an email when Qwen has news
No news that day, no email.
More stories today
- Claude Code 2.1.251 adds model-switch hooks, subagent streaming
- LM Studio's AI command judge starts agreeing with the defendant
- AMD releases ROCm 10.0 with native agentic AI developer experience
- UnifyGTM cuts agent model costs by 90-95%
- Criticism of AI hype: 'magical machine god' argument distracts from real harms