AppleAnalysisAI ModelsSeptember 16, 2026

Apple researchers propose DACA-GRPO for diffusion language models

DACA-GRPO adds Denoising Progress Scores and Stratified Masking Likelihood to any GRPO-style trainer, targeting missing temporal credit assignment and mean-field likelihood bias. Gains reach 36.3pp on constraint satisfaction, 7.4pp on code generation, and 5.6pp on math reasoning across seven benchmarks.

1 source

More stories today

Open the live feed