Alibaba releases Qwen3.8-2.4T-A95B open weights; NVIDIA serves it on GB300 NVL72

Qwen3.8-Max is a sparse MoE with 2.4T total parameters, 95B active per token, and a 1M-token context. NVIDIA reports Day-0 throughput above 4K tokens/s per GPU on GB300 NVL72 in FP8, with Qwen calling it its first multimodal model above 1T parameters.
How this story unfolded
6 days · 2 reports · 5 community posts · 7 of 8 shown
- Aug 12
- Aug 13
- Aug 14
- Aug 18
Qwen by email
Get an email when Qwen has news
No news that day, no email.
More stories today
- Avvoka Partners With Harvey, Launches Curate For Templates
- How to audit preference biases and fine-tune with DPO (TRL + LoRA)
- CLAUDE.md applies Karpathy's engineering principles to Claude Code
- Hays Shifts to Hard-to-Replace Roles as AI Reshapes Hiring
- Claude says I used 54.9 BILLION tokens.