AnalysisDevelopersJuly 20, 2026
NInfer inference engine hits 543 tok/s for Qwen3.6-35B-A3B on RTX 5090
NInfer is a from-scratch C++/CUDA engine, now open-sourced, specialized for two exact Qwen3.6 checkpoints on a single RTX 5090; the 543 tok/s run spans a 65K-token decode. Code is on GitHub at github.com/Neroued/ninfer.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Thoughtworks' Kief Morris: humans must stay 'on the loop' in AI delivery
- GEMA wins major copyright ruling against Suno, orders damages paid
- LangChain builds ReviewBench benchmark for code review agents
- DeepSeek Flash 0731's reasoning trace amuses with 'OH MY GOD' outburst
- Former OpenAI VP Jerry Tworek discusses AI lab automation