AnalysisAI ModelsJune 24, 2026
DFlash speculative decoding achieves up to 15x throughput on NVIDIA Blackwell

DFlash drafts whole token blocks in parallel, overcoming autoregressive bottleneck. Claims up to 15x higher throughput on NVIDIA Blackwell GPUs, especially for long chain-of-thought reasoning models.