AnalysisAI ModelsOctober 4, 2026

Qwen3.8-Flash-Next 125B runs on consumer hardware via Strata and Kyojin engines

Read original source →github.com

Community builders got the 125B MoE (6B active) model running on single consumer boxes: 44-59 tok/s on an AMD Strix Halo mini PC with 95 GB EXL3 weights, and 150-200 tok/s decode on a power-limited RTX 5090 with 96GB DDR5. A 16GB RTX 3080 laptop and 12GB RTX 5070 also run it via TensorSharp and llama.cpp.

How this story unfolded

3 weeks · 1 report · 23 community posts · from Sep 15

  1. Sep 15
  2. Sep 16
  3. Sep 18
  4. Sep 19
  5. Sep 20
  6. Sep 24
  7. Oct 1
  8. Oct 2
  9. Oct 3
  10. Oct 4
  11. Oct 5

More stories today

Open the live feed