Alibaba releases Qwen3.8-2.4T-A95B, its largest open-weight model

Qwen3.8-2.4T-A95B (Qwen3.8-Max) has 2.4T total parameters with 95B activated per token, a 1M-token context window, and 128K output length. It achieves over 4K tokens/s per GPU on NVIDIA GB300 NVL72 in FP8 on Day 0.
How this story unfolded
1 day · 3 reports · 2 community posts · 5 of 10 shown
- Aug 12
- Aug 13
Qwen by email
Get an email when Qwen has news
No news that day, no email.
More stories today
- Skeptic shares first impressions of ChatGPT Plus
- Hard sci-fi authors largely oppose LLMs, survey finds
- Anthropic tests folderless Claude Code sessions on Desktop and iOS
- Ethan Mollick: Using weaker AI for human-facing content may soon be disrespectful
- Pocket TTS training stack open-sourced