AnalysisAI ModelsOctober 3, 2026

Qwen3.8-Flash-Next 125B MoE runs locally on Strix Halo at 44-59 tok/s

Read original source →github.com

Qwen3.8-Flash-Next (125B MoE, 6B active) runs on a single AMD Strix Halo mini PC (Ryzen AI Max+ 395, 128GB) at 44-59 tok/s with speculative decoding and ~1,400 tok/s prefill. The 95GB EXL3 weights and Kyojin inference engine (built on ExLlamaV3) were released alongside.

How this story unfolded

5 days · 0 reports · 11 community posts · 11 of 12 shown

  1. Oct 1
  2. Oct 2
  3. Oct 3
  4. Oct 5
  5. Oct 6

More stories today

Open the live feed