AnalysisAI ModelsAugust 16, 2026

Qwen 3.8 2.4T hits 288k tokens/s on Nvidia GB300 NVL72

Nvidia's blog reports serving the 2.4T-parameter Qwen 3.8 model at 288k tokens/s on GB300 NVL72, with 4k tokens/s per GPU.

1 source

AI Models by email

Get an email when there's news on AI Models

No news that day, no email.

More stories today

Open the live feed