Spatial gradient maps improve RL alignment of flow-based image models

The authors introduce a token-wise credit assignment method for Diffusion Transformers that uses reward gradients to assign localized advantages, achieving state-of-the-art alignment on GenEval.

Chinese Tech
Yunkai Yang · Yudong Zhang · Xinying Chen · Bin Luo · Jienan Lyu · Kunquan Zhang · +2 more

Sun Yat-Sen University · Baidu · Beijing Institute of Technology · Tsinghua University · TS Martech

Research Digest··2 min read
The paper proposes SGA-Flow-GRPO, a reinforcement learning framework that assigns spatially differentiated credit to individual latent tokens in flow-based generative models.

The authors modified the Flow-GRPO reinforcement learning algorithm to account for the spatial structure of images generated by Diffusion Transformers (DiTs).

Why this paper

From Baidu and 4 others

In one line

SGA-Flow-GRPO uses spatial gradient-guided credit assignment to achieve state-of-the-art alignment quality for flow-based generative models on GenEval.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.