AnalysisAI ModelsSeptember 17, 2026

Qwen3.6 35B-A3B hits 600 tok/s on RTX Pro 6000 with Ninfer

A single request on Qwen3.6 35B-A3B reached 600 tokens/sec running Ninfer on an RTX Pro 6000. The poster calls it a "drudgework" model for read-and-find or code tasks, noting it can burn 20x more tokens and still beat many local models.

1 source

More stories today

Open the live feed