NVIDIALaunchAI ModelsAugust 12, 2026

Alibaba releases Qwen3.8-Max, a 2.4T-parameter open-weight model

Qwen3.8-Max has 2.4T total parameters with 95B activated per token, a fine-grained MoE with hybrid full/linear attention and up to 1M-token context. On NVIDIA GB300 NVL72 it serves over 4K tokens/s per GPU in FP8 on day one; Together AI already hosts it.

How this story unfolded

6 days · 1 report · 5 community posts · from Aug 12

  1. Aug 12
  2. Aug 13
  3. Aug 14
  4. Aug 18

NVIDIA by email

Get an email when NVIDIA has news

No news that day, no email.

More stories today

Open the live feed