AnalysisAI ModelsSeptember 24, 2026

Custom engine runs Qwen3.8-Flash-Next at 65 tok/s on 12GB VRAM

Read original source →reddit.com

A developer's custom inference engine pushes the IQ3_XXS quant of Qwen3.8-Flash-Next to ~65 tok/s output and ~430 tok/s prompt processing on a 12GB RTX 5070, up from 15 tok/s output and 100-120 tok/s prompt processing in llama.cpp.

1 source

More stories today

Open the live feed