The authors modified the Flow-GRPO reinforcement learning algorithm to account for the spatial structure of images generated by Diffusion Transformers (DiTs).
Spatial gradient maps improve RL alignment of flow-based image models
The authors introduce a token-wise credit assignment method for Diffusion Transformers that uses reward gradients to assign localized advantages, achieving state-of-the-art alignment on GenEval.
Chinese Tech
Yunkai Yang · Yudong Zhang · Xinying Chen · Bin Luo · Jienan Lyu · Kunquan Zhang · +2 more
Sun Yat-Sen University · Baidu · Beijing Institute of Technology · Tsinghua University · TS Martech
Research Digest··2 min read
The paper proposes SGA-Flow-GRPO, a reinforcement learning framework that assigns spatially differentiated credit to individual latent tokens in flow-based generative models.
Why this paper
From Baidu and 4 others
In one line
SGA-Flow-GRPO uses spatial gradient-guided credit assignment to achieve state-of-the-art alignment quality for flow-based generative models on GenEval.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§