AnalysisDevelopersAugust 21, 2026

Qwen3.8-27B hits 381 tps on a single RTX 3090

A hyper-optimized Qwen3.8-27B inference engine reached 381 tps for a single request on an RTX 3090, up from 82 tps at release four days earlier. The 3090 is power limited to 250W; peak throughput was 672 tps.

How this story unfolded

2 days · 0 reports · 3 community posts · from Aug 18

  1. Aug 18
  2. Aug 19
  3. Aug 21

More stories today

Open the live feed