Human critiques help generative reward models produce more reliable evaluations

EnGRICH learns evaluation criteria from limited human critiques, then applies them to preference data that contains only outcome labels.

Chinese Tech
Xuancheng Li · Beining Wang · Haitao Li · Heng Wang · Yujia Zhou · Qingyi Pan · +4 more

Tsinghua University · Tencent

Research Digest··2 min read
Li and colleagues address a weakness in generative reward models: choosing the preferred response correctly does not guarantee that the accompanying critique is sound.

The authors trained a generative reward model, or GRM, to compare two responses, explain its reasoning in natural language and select the better one.

Why this paper

From Tencent and Tsinghua University

In one line

EnGRICH uses human critiques to train a MetaCritic that improves generative reward model training with process supervision and guided exploration.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.