DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models
This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.
Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood estimates. We identify two fundamental weaknesses: the absence of temporal credit…
Read the full story at Apple Machine Learning Research ↗
Timeline · 1 report
- 2026-09-16 00:00 · Apple Machine Learning Research
DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models