New papers and llama.cpp support advance speculative decoding

Four arXiv papers propose speculative decoding methods: TreeGraft, AgentSpec, LiLiCorr, and Self-Speculation. Meanwhile, llama.cpp merged support for DFlash2, a diffusion-style block head drafter.
How this story unfolded
4 days · 4 reports · 1 community post · from Aug 24
- Aug 24
- Aug 26
- Aug 27
- Aug 28
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Claude Code silently deletes local history older than 30 days
- Split sigmas speed up video reference in Stable Diffusion
- MiniMax tool turns anything into realistic human video
- MirroS' Code-as-World rewrites videos into executable MuJoCo programs
- PRAXIST research system boosts MLE-bench scores 44%