AnalysisAI ModelsSeptember 3, 2026

Qwen3.8-Flash-Next hits 25-29 t/s on 2x3090 with expert cache PR

A dual RTX 3090 (PCIe 3.0) + dual Xeon E5-2696 v4 setup with 188 GB DDR4-2133 LRDIMM running llama.cpp and an unsloth UD-Q6_K_XL quant went from 17 to 25-29 t/s decode after applying the expert cache PR, with all 48 expert layers pinned in host memory.

1 source

More stories today

Open the live feed