AnalysisAI ModelsOctober 9, 2026

Qwen 3.8 Flash Next runs at 21 tok/s on RTX 3060 12GB

Read original source →reddit.com

A community quantization (GSQ-RCO-IQ2_XS) of Qwen 3.8 Flash Next hits ~21 tok/s on an RTX 3060 12GB plus 16GB DDR4, rising to 24+ tok/s with a warm cache. The poster reports no gate pruning and 100% bit-exact output, following an earlier experiment on predicting MoE expert usage for CPU/GPU offloading.

1 source

More stories today

Open the live feed