AnalysisAI ModelsSeptember 8, 2026

GSQ-RCO GGUFs released for Qwen3.8-Flash-Next, plus 50% expert-pruned build

Read original source →huggingface.co

Qwen3.8-Flash-Next is a sparse MoE with 512 routed experts per layer across 48 layers and 176.9B parameters, 354 GB unquantized. The GSQ-RCO quants claim near-baseline performance, and a second capability-targeted build drops half the experts at ~1.89 bpw.

How this story unfolded

13 days · 1 report · 2 community posts · from Sep 16

  1. Sep 16
  2. Sep 17
  3. Sep 29

More stories today

Open the live feed