Dflash2: Qwen3.8-27B hits 138 tps on RTX 3090
A hyper-optimized inference engine for Qwen3.8-27B on an RTX 3090 reaches 138 tokens per second (power-limited to 250W), up from 82 tps single request in the prior release.
1 source
AI Models by email
Get an email when there's news on AI Models
No news that day, no email.
More stories today
- Lyte closes $165M round at $1.6B valuation
- Meta settlement could clear way for new AI product launches
- Z.ai opens first Tmall store for AI subscriptions
- Fable 5.1 Max users share setup tips and warnings
- Opinion: Next DSM should assess algorithms' role in eating disorders