Within each group of images generated for a prompt, MEND caps rewards at a chosen quantile.
MEND selectively moves flow-model samples when reward gains justify distance
The method improves differentiable rewards by accepting per-sample updates only when their gains exceed a quadratic movement cost.
Big Tech
Shreshth Saini · Neil Birkbeck · Yilin Wang · Balu Adsumilli · Alan C. Bovik
The University of Texas at Austin · Google · University of Colorado Boulder
Research Digest··2 min read
Saini and colleagues introduce MEND, a reinforcement learning method for post-training text-to-image flow models without KL penalties, frozen reference models or advantage weighting.
Why this paper
From Google and 2 others
In one line
MEND caps rewards per group and accepts a flow-model velocity update only when the capped reward gain exceeds a quadratic displacement price.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§