AnalysisAI ModelsAugust 23, 2026

Qwen3.8-27B NVFP4 runs on RTX 5090 at 120 tok/s

A user reports running Qwen3.8-27B NVFP4 with vision and 451K-token KV-cache on a single RTX 5090 (power-limited to 400W), averaging 120 tokens/s. The setup targets 3 parallel requests with concurrency enabled.

1 source

More stories today

Open the live feed