AnalysisDevelopersSeptember 7, 2026

ExllamaV3 beats llama.cpp on CPU-offloaded Qwen 3.8 Flash Next

On 2x 20GB RTX 3080s, 128GB DDR4 and a Xeon 6148, llama.cpp ran Unsloth's Q4_K_XL quant of Qwen 3.8 Flash Next at ~270tps prefill and ~13tps decode, dropping from near 20tps.

1 source

More stories today

Open the live feed