AnalysisDevelopersJuly 27, 2026

Nifer inference engine hits 700 t/s with Qwen 3.6 35B on RTX 5090

A user running Nifer on Windows reports 550-720 t/s single-instance with Qwen 3.6 35B (no thinking) on an RTX 5090, including full 250k context. The poster says matching speeds previously required batching and parallel agents.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed
Nifer inference engine hits 700 t/s with Qwen 3.6 35B on RTX 5090 — AIBriefs