AnalysisAI ModelsOctober 8, 2026

Qwen3.8-Flash-Next runs locally on Strix Halo via Strata and Kyojin

Read original source →reddit.com

Qwen3.8-Flash-Next (125B MoE, 6B active, plus 51B n-gram embedding and 4B MTP) hits 44-59 tok/s decode and ~1,400 tok/s prefill on a single AMD Strix Halo mini PC with speculative decoding. Strata adds official Strix Halo support with up to 1M context without big speed loss.

How this story unfolded

4 days · 0 reports · 5 community posts · from Oct 2

  1. Oct 2
  2. Oct 3
  3. Oct 5
  4. Oct 6

More stories today

Open the live feed