AnalysisAI ModelsAugust 31, 2026

Qwen3.8-Flash-Next hits 120 t/s on 4xR9700 with vLLM

A Reddit user reports running Qwen3.8-Flash-Next on 4xR9700 GPUs at 80-120 tokens/s generation and 12k tokens/s prefill for a single request, using an optimized vLLM setup.

1 source

AI Models by email

Get an email when there's news on AI Models

No news that day, no email.

More stories today

Open the live feed