AnalysisAI ModelsSeptember 14, 2026

Qwen3.8-Flash-Next runs locally on 12-16GB GPUs via expert streaming

Read original source →carteakey.dev

Community builds (Strata, TensorSharp, llama.cpp forks) run the 125B-A6B MoE plus 51B n-gram table on 12-16GB cards, since the n-gram table needs no RAM. Reported throughput ranges from 11.5 tok/s on an RTX 5070 to 65 tok/s on the same card after a custom engine.

How this story unfolded

3 weeks · 0 reports · 15 community posts · from Sep 14

  1. Sep 14
  2. Sep 15
  3. Sep 16
  4. Sep 24
  5. Sep 30
  6. Oct 1
  7. Oct 2
  8. Oct 3
  9. Oct 4
  10. Oct 6

More stories today

Open the live feed