AnalysisAI ModelsSeptember 4, 2026

Qwen3.8-Flash-Next hits 37-41 t/s on 2x3090 with expert cache

A user reports 37-41 t/s decode on 2x RTX 3090 with DDR4, up from 25-29 t/s, using UD-Q4_K_XL quantization, expert cache, and MTP in llama.cpp. The setup keeps all 48 expert layers in host RAM with full 261k context.

1 source

More stories today

Open the live feed