AnalysisDevelopersAugust 19, 2026

Dflash2 engine hits 138 tps on RTX 3090

A hyper-optimized Qwen3.8-27B inference engine for RTX 3090 reaches 138 tokens per second single request, up from 82 tps in the previous release. The GPU is power-limited to 250W.

How this story unfolded

3 weeks · 0 reports · 14 community posts · from Aug 14

  1. Aug 14
  2. Aug 16
  3. Aug 18
  4. Aug 19
  5. Aug 21
  6. Aug 23
  7. Aug 29
  8. Aug 31
  9. Sep 1
  10. Sep 2
  11. Sep 3
  12. Sep 5

More stories today

Open the live feed