AnalysisAI ModelsAugust 16, 2026

Qwen3.8-27B hits 82 tps on RTX 3090

A custom inference engine achieves 82 tps single request and up to 672 tps peak on a 250W power-capped RTX 3090, with 417 tps sustained at 64 concurrent requests. Supports up to 195k context (150k shipped).

1 source

AI Models by email

Get an email when there's news on AI Models

No news that day, no email.

More stories today

Open the live feed
Qwen3.8-27B hits 82 tps on RTX 3090 — AIBriefs