LaunchDevelopersJuly 13, 2026

DFlash speeds up Qwen models 2.2x in llama.cpp

First benchmarks show DFlash achieves 98 tok/s on Qwen3.6-27B, a 2.2x speedup over baseline with byte-for-byte identical outputs. The technique uses a small draft model to generate up to 15 tokens at once, with the full model only verifying.

5 sources