New papers advance diffusion LLM decoding and speculative decoding
Multiple arXiv papers propose methods to speed up diffusion language models and speculative decoding, including visual-information-guided parallel decoding, survival-guided length control, and adaptive draft-tree construction. Techniques target efficient inference for masked diffusion models and LLM agents.
How this story unfolded
4 days · 12 reports · from Aug 24
- Aug 24
Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosisarxiv.org
Self-Speculation for Faster Reasoning Modelsarxiv.org
GRAFT: Adaptive DLM-Based Draft Tree Construction with Target-Distilled Edge Scoringarxiv.org
LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decodingarxiv.org
- Aug 26
- Aug 27
- Aug 28
Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Modelsarxiv.org
TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decodingarxiv.org
Affix Cache for Diffusion Large Language Modelsarxiv.org
Survival-Guided Length Control for Efficient Diffusion Language Modelsarxiv.org
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- SentinelOne CEO discusses earnings and AI's cybersecurity impact
- Andrew Ng: Biggest AI opportunities aren't where you think
- OpenAI co-founder warns of closing window to secure internet
- Yann LeCun: Provably safe AI is impossible
- Gemini 3.8 Flash preview reportedly in use at Google