AnalysisDevelopersJuly 20, 2026

NInfer inference engine hits 543 tok/s for Qwen3.6-35B-A3B on RTX 5090

NInfer is a from-scratch C++/CUDA engine, now open-sourced, specialized for two exact Qwen3.6 checkpoints on a single RTX 5090; the 543 tok/s run spans a 65K-token decode. Code is on GitHub at github.com/Neroued/ninfer.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed
NInfer inference engine hits 543 tok/s for Qwen3.6-35B-A3B on RTX 5090 — AIBriefs