NVIDIA serves Qwen3.8-2.4T-A95B on GB300 NVL72

Alibaba released open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), a 2.4T-parameter MoE model with 95B activated per token. On NVIDIA GB300 NVL72, it achieves over 4K tokens/s per GPU and 350 tokens/s per user in FP8 on Day 0.
1 source
NVIDIA by email
Get an email when NVIDIA has news
No news that day, no email.
More stories today
- Claude Code 2.1.251 adds model-switch hooks, subagent streaming
- LM Studio's AI command judge starts agreeing with the defendant
- AMD releases ROCm 10.0 with native agentic AI developer experience
- UnifyGTM cuts agent model costs by 90-95%
- Criticism of AI hype: 'magical machine god' argument distracts from real harms