AnalysisDevelopersJuly 16, 2026

llama.cpp speculative decoding: n-gram stack ~6x on Qwen 3.6 27B

On an RTX 6000 PRO with Qwen 3.6 27B, llama.cpp's n-gram stack hit ~6x speedup, DFlash ~3.7x, and MTP ~2.7x on real coding. DFlash recently merged via PR #22105; an earlier run posted 4.44x at 36K context.

2 sources

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed