AnalysisAI ModelsJune 30, 2026

User benchmarks Qwen 3.6 27B at ~100 TPS on RTX 3090

A Reddit user reports ~100 tokens per second on Qwen 3.6 27B using speculative decoding on a single RTX 3090, comparing 5 inference engines including llama.cpp forks. The benchmark highlights efficient local inference achievable with consumer hardware.

1 source

More stories today

Open the live feed