LaunchAI ModelsJune 23, 2026

Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative…

DFlash, an open-source block diffusion model for speculative decoding, achieves up to 15x inference performance for gpt-oss-120b on NVIDIA Blackwell and nearly doubles interactivity for Llama 3.1 8B vs EAGLE-3. The research team has released 20 DFlash checkpoints on Hugging Face.

1 source