QwenLaunchAI ModelsAugust 12, 2026

Alibaba releases Qwen3.8-2.4T-A95B open-weight model

Alibaba released Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model with 2.4T total parameters and 95B activated per token. It features a fine-grained MoE architecture with hybrid full and linear attention, a 1M-token context window, and up to 128K output tokens. On NVIDIA GB300 NVL72, it achieves over 4K tokens per second per GPU in FP8 precision.

How this story unfolded

4 days · 4 reports · 6 community posts · 10 of 15 shown

  1. Aug 12
  2. Aug 13
  3. Aug 14
  4. Aug 16

More stories today

Open the live feed