LaunchDevelopersAugust 18, 2026

Inco AI releases DFlash 2 parallel speculative decoding

DFlash 2 boosts output per verification pass by over 20% with ~1% added latency, gains 16–25% across benchmarks. SGLang with the new Qwen3.8-27B drafter serves at 2.7–3.4× autoregressive throughput at batch size 1.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed