LaunchDevelopersAugust 18, 2026

Inco AI unveils DFlash 2 parallel speculative decoding

DFlash 2 boosts output per verification pass by over 20% with ~1% added latency, gains 16–25% across benchmarks. With the new Qwen3.8-27B drafter, SGLang serves at 2.7–3.4× autoregressive throughput at batch size 1.

1 source

Developers by email

Get an email when there's news on Developers

No news that day, no email.

More stories today

Open the live feed
Inco AI unveils DFlash 2 parallel speculative decoding