llama.cpp speculative decoding: n-gram stack ~6x on Qwen 3.6 27B

On an RTX 6000 PRO with Qwen 3.6 27B, llama.cpp's n-gram stack hit ~6x speedup, DFlash ~3.7x, and MTP ~2.7x on real coding. DFlash recently merged via PR #22105; an earlier run posted 4.44x at 36K context.
2 sources
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Post-training course materials invite educator feedback
- Kimi K3 available to try free on Together Chat
- Rhodium's Goujon urges holistic AI safety approach
- Cheap AI intelligence revives graph knowledge and ontologies
- US will exempt Chinese open-weight models from safety testing requirements