AnalysisDevelopersSeptember 1, 2026

Developer pushes Qwen3.8-27B to 2,000 prefill tps on RTX 3090

A custom inference engine for Qwen3.8-27B now hits 2,000 prefill tokens/sec and 132 decode tokens/sec on a single power-limited RTX 3090. The same build reached 381 tps single-request earlier, up from 82 tps in the initial release.

1 source

More stories today

Open the live feed