AnalysisAI ModelsAugust 23, 2026

Qwen3.8-27B NVFP4 runs with vision and 451K KV-cache on RTX 5090

A user reports running Qwen3.8-27B NVFP4 with vision and a 451K-token KV-cache on a single RTX 5090 power-limited to 400W, averaging 120 tokens/s. The setup targets 3 parallel requests with concurrency enabled.

1 source

More stories today

Open the live feed
Qwen3.8-27B NVFP4 runs with vision and 451K KV-cache on RTX 5090