AnalysisDevelopersSeptember 13, 2026

Reddit user benchmarks Qwen3.8 27B INT4 on RTX 3090 via vLLM

vLLM ran Qwen3.8 27B INT4 with 144K FP8 KV cache on a 24GB RTX 3090, averaging 871.93 tok/s prefill and 38.39 tok/s decode over 3 hours. A 10K prompt hit 1000.26 tok/s prefill and 42.3 tok/s decode.

How this story unfolded

8 days · 0 reports · 6 community posts · from Sep 5

  1. Sep 5
  2. Sep 6
  3. Sep 7
  4. Sep 12
  5. Sep 13

More stories today

Open the live feed