Diffusion reward models capture multimodal preference that scalar scores miss

A lightweight Diffusion Transformer conditioned on a frozen LLM encoder estimates full reward distributions without parametric assumptions, matching or beating baseline reward models on five benchmarks and improving downstream RLHF.

Top University
Xiangyang Wang · Bingxiang He · Zeyuan Liu · Jiaze WangZiqing Qiao · Yuxin Zuo · Huan-ang Gao · +9 more

Tsinghua University · The Chinese University of Hong Kong · University of Illinois Urbana-Champaign

Research Digest··3 min read
Wang et al.

The authors recast reward modeling as conditional density estimation over p(\mathbf{r}|x,y).

Why this paper

From Tsinghua University and 2 others · Part of RL for Tool Agents, now 13 papers

In one line

DRM models reward as a multimodal conditional density, matching or surpassing baseline reward models.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.