AnalysisAI ModelsOctober 3, 2026

Qwen3.8-Flash-Next 177B runs at 11-15 tok/s on single RTX 5070 12GB

Read original source →github.com

A llama.cpp expert-streaming setup on Windows pushes Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) to about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup, using a 12GB RTX 5070 plus 32GB DDR4 RAM.

1 source

More stories today

Open the live feed