Qwen3.8-Flash-Next hits 120 t/s on 4xR9700 with vLLM

A Reddit user reports running Qwen3.8-Flash-Next on 4xR9700 GPUs at 80-120 tokens/s generation and 12k tokens/s prefill for a single request, using an optimized vLLM setup.
1 source
AI Models by email
Get an email when there's news on AI Models
No news that day, no email.
More stories today
- Open-source RL training with trl and OpenEnv shared
- NVIDIA invests $3.5B in MediaTek, deepens AI partnership
- Neta team explains why their open-source model generated Anne Hathaway-like images
- OpenAI age-verification error deletes adult's account
- South Korea gives citizens free unlimited domestic AI access