The authors recast reward modeling as conditional density estimation over p(\mathbf{r}|x,y).
Diffusion reward models capture multimodal preference that scalar scores miss
A lightweight Diffusion Transformer conditioned on a frozen LLM encoder estimates full reward distributions without parametric assumptions, matching or beating baseline reward models on five benchmarks and improving downstream RLHF.
Top University
Xiangyang Wang · Bingxiang He · Zeyuan Liu · Jiaze WangZiqing Qiao · Yuxin Zuo · Huan-ang Gao · +9 more
Tsinghua University · The Chinese University of Hong Kong · University of Illinois Urbana-Champaign
Research Digest··3 min read
Thread:RL for Tool Agents
Wang et al.
Why this paper
From Tsinghua University and 2 others · Part of RL for Tool Agents, now 13 papers
In one line
DRM models reward as a multimodal conditional density, matching or surpassing baseline reward models.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§