New arXiv papers target faster diffusion language model inference

Papers propose cached hidden-state reuse (Archer), linear attention retrofits, and draft-then-refine decoding to cut rollout cost for diffusion language models, which iteratively refine text instead of generating left-to-right. Separate work speeds speculative decoding for MoE targets and low-latency LLM agent tool calls.
2 sources
More stories today
Pinloop CLI lets coding agents run job searches
Vaibhav Sisinty·3 hours ago
Minimax H3 generates Sailor Moon clips shared on r/StableDiffusion
Two r/StableDiffusion posts by the same user showcase Minimax H3 video generations of Sailor Moon characters, including a Mina clip and a Sailor Jupiter clip. The poster noted unintended zoom-ins in the Mina output.
r/StableDiffusion·3 hours ago
Suno V6 users report audio quality degradation late in songs
Reddit users on r/SunoAI report severe audio quality degradation toward the end of Suno V6 tracks, including when Max Mode is enabled. The poster says structure, rhythm, instruments, lyrics, and vocals stay consistent, so the issue is not the song composition itself.
r/SunoAI·4 hours agoClaude Opus 5 controls SO-101 robot arm to paint
A Reddit user gave Claude Opus 5 control of an SO-101 robot arm fitted with a paintbrush and asked it to paint the Golden Gate Bridge. The model learned the arm's controls, wrote calibration files, and broke the painting task into simpler instructions, producing progressively better paintings.
r/ClaudeAI·4 hours ago
Lemonade 2026.40 RC fixes AMD APU model streaming, drops OpenMOSS ROCm
Lemonade's 2026.40 release candidate sizes streaming models against the APU GTT pool instead of the fixed vRAM carve-out, fixing failures like DeepSeek-V4-Flash-IQ2XXS-DS4 on Ryzen AI Max (Strix Halo). OpenMOSS ROCm back-ends were removed on Windows and Linux after running ~40x slower than Vulkan.
r/LocalLLaMA·4 hours ago
Reddit user runs Qwen FN on Strix box, compares it to 27B
A LocalLLaMA user reports running Qwen FN on a Strix box for several days, crediting the Halogen team for performance they call "the setup" for that hardware and model.
r/LocalLLaMA·4 hours agoChina's AI chip blitz arms Xi with a message for Trump
A run of new chips and AI models from Chinese national champions gives Xi Jinping a confidence boost as the U.S. summit gets underway, CNBC reports. The piece frames the export-control standoff as "you can't choke us off."
CNBC Technology·4 hours ago

San Jose Sharks drop Suno-generated goal song after fan backlash
The NHL team replaced the AI-generated goal song after it lasted just one season, following fan discovery that it was made with Suno. A new goal song is being introduced.
Digital Music News·4 hours ago
