AnalysisAI ModelsJuly 7, 2026
DFlash in llama.cpp yields 4.44x speedup on Qwen 3.6 27B

DFlash (block diffusion speculative decoding) merged into llama.cpp achieves 4.44x faster generation at 36K context on an RTX 6000 PRO with Qwen 3.6 27B. This beats previous MTP benchmarks by a wide margin across all draft lengths tested.