AnalysisAI ModelsOctober 3, 2026

Qwen3.8-Flash-Next 177B runs at 11-15 tok/s on single RTX 5070 12GB

Read original source →github.com

A llama.cpp expert-streaming setup on Windows hits ~11.5 tok/s with the UD-IQ3_XXS quant of Qwen3.8-Flash-Next 177B, up from roughly 7 tok/s on the inherited setup. Hardware is an RTX 5070 with 12GB VRAM plus 32GB DDR4 RAM.

1 source

More stories today

Open the live feed