AnalysisDevelopersSeptember 13, 2026

Reddit user benchmarks Qwen3.8 27B INT4 on RTX 3090 via vLLM

A 24GB RTX 3090 run of Qwen3.8 27B INT4 with 144K FP8 KV cache averaged 871.93 tok/s prefill and 38.39 tok/s decode over 3 hours of constant running. A 10K prompt hit 1000.26 tok/s prefill and 42.3 tok/s decode.

1 source

More stories today

Open the live feed