AnalysisAI ModelsJune 30, 2026
User benchmarks Qwen 3.6 27B at ~100 TPS on RTX 3090
A Reddit user reports ~100 tokens per second on Qwen 3.6 27B using speculative decoding on a single RTX 3090, comparing 5 inference engines including llama.cpp forks. The benchmark highlights efficient local inference achievable with consumer hardware.
1 source
More stories today
- Visa open-sources Mythos harness for payment network bug hunting
- Codeberg votes to ban AI-written code projects
- En route to improving your agents
- NVIDIA GTC SJ 2026: AI-native digital health stack guide
- Future AI acceleration may need pacing, says OpenAI