AnalysisAI ModelsAugust 16, 2026

Qwen3.8-27B hits 82 tps on RTX 3090

A Reddit user reports a custom inference engine for Qwen3.8-27B on an RTX 3090, achieving 82 tokens/s single request and up to 672 tps peak. Supports up to 195k context, ships with 150k for safety.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed