Qwen 3.8 2.4T hits 288k tokens/s on Nvidia GB300 NVL72

Nvidia's blog reports serving the 2.4T-parameter Qwen 3.8 model at 288k tokens/s on GB300 NVL72, with 4k tokens/s per GPU.
1 source
AI Models by email
Get an email when there's news on AI Models
No news that day, no email.
More stories today
- Open-source RL training with trl and OpenEnv shared
- NVIDIA invests $3.5B in MediaTek, deepens AI partnership
- Neta team explains why their open-source model generated Anne Hathaway-like images
- OpenAI age-verification error deletes adult's account
- South Korea gives citizens free unlimited domestic AI access