Alibaba releases Qwen3.8-Max, a 2.4T-parameter open-weight model

Qwen3.8-Max has 2.4T total parameters with 95B activated per token, a fine-grained MoE with hybrid full/linear attention and up to 1M-token context. On NVIDIA GB300 NVL72 it serves over 4K tokens/s per GPU in FP8 on day one; Together AI already hosts it.
How this story unfolded
6 days · 1 report · 5 community posts · from Aug 12
- Aug 12
- Aug 13
- Aug 14
- Aug 18
NVIDIA by email
Get an email when NVIDIA has news
No news that day, no email.
More stories today
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills
- Aident Loadout gives agents 27,000+ tools and logs every action
- Etched gains sizable fan base for AI inferencing computers
- Corbell generates technical specs from repository knowledge graphs