Qwen3.8-27B hits 82 tps on RTX 3090

A Reddit user reports a custom inference engine for Qwen3.8-27B on an RTX 3090, achieving 82 tokens/s single request and up to 672 tps peak. Supports up to 195k context, ships with 150k for safety.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Redditor tests AI agents with $1 online task
- AI agents need their own identity before a gateway
- Claude Max users find default $200K spend limit
- TTFT-First Benchmark Ranks Lowest-Latency Voice and Realtime Agent APIs
- AI training demand causes Mac Mini shortages