AnalysisAI ModelsSeptember 27, 2026

GSQ-RCO GGUFs released for Qwen3.8-Flash-Next, plus 50% expert-pruned Coder build

Read original source →huggingface.co

ISTA-DASLab published GSQ and RCO quantized GGUFs of Qwen3.8-Flash-Next, a sparse MoE with 512 routed experts per layer across 48 layers and 176.9B parameters (354 GB unquantized). A second capability-targeted build strips half the experts at ~1.89 bpw.

How this story unfolded

2 days · 1 report · 2 community posts · from Sep 28

  1. Sep 28
  2. Sep 29
  3. Oct 1

More stories today

Open the live feed