AnalysisAI ModelsAugust 13, 2026

New arXiv papers target faster diffusion language model inference

Papers propose cached hidden-state reuse (Archer), linear attention retrofits, and draft-then-refine decoding to cut rollout cost for diffusion language models, which iteratively refine text instead of generating left-to-right. Separate work speeds speculative decoding for MoE targets and low-latency LLM agent tool calls.

2 sources

More stories today

Open the live feed