The authors formalized the optimization objective underlying regression-based diffusion RL: maximize reward on a frozen rollout buffer under a divergence penalty.
Three diffusion RL methods unified as divergence-constrained reward maximizers
Authors show DiffusionNFT, FlowAWR, and RAM share a common framework; their new method DiffusionRFT achieves faster convergence and stability.
Chinese Tech
Toyota Li · David Zhao · Alan Zhao
Tencent
Research Digest··3 min read
A unified theoretical framework reveals that three prominent regression-based diffusion reinforcement learning methods—DiffusionNFT, FlowAWR, and RAM—are each solutions to a divergence-constrained reward-maximization problem, differing only in the convex generator defining the constraint.
Why this paper
From Tencent
In one line
DiffusionRFT uses exact sparsemax projection to converge faster, train more stably, and achieve top performance in regression-based diffusion RL.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§