Qwen3.8-27B hits 82 tps on RTX 3090

A custom inference engine achieves 82 tps single request and up to 672 tps peak on a 250W power-capped RTX 3090, with 417 tps sustained at 64 concurrent requests. Supports up to 195k context (150k shipped).
1 source
AI Models by email
Get an email when there's news on AI Models
No news that day, no email.
More stories today
- Lyte closes $165M round at $1.6B valuation
- Meta settlement could clear way for new AI product launches
- Z.ai opens first Tmall store for AI subscriptions
- Fable 5.1 Max users share setup tips and warnings
- Opinion: Next DSM should assess algorithms' role in eating disorders