AnalysisAI ModelsOctober 3, 2026

LocalLLaMA users push Qwen3.8-Flash-Next onto consumer GPUs

Read original source →github.com

Community benchmarks run the 176-180B MoE model (125B params, 6B activated, plus 51B n-gram embedding and 4B MTP) on 12-16GB cards, hitting 65 tok/s on an RTX 5070 and 150-200 tok/s on a 5090 with 96GB DDR5.

How this story unfolded

3 weeks · 0 reports · 20 community posts · from Sep 15

  1. Sep 15
  2. Sep 16
  3. Sep 18
  4. Sep 19
  5. Sep 20
  6. Sep 24
  7. Sep 28
  8. Oct 1
  9. Oct 2
  10. Oct 3
  11. Oct 4

More stories today

Open the live feed