The authors trained a generative reward model, or GRM, to compare two responses, explain its reasoning in natural language and select the better one.
Human critiques help generative reward models produce more reliable evaluations
EnGRICH learns evaluation criteria from limited human critiques, then applies them to preference data that contains only outcome labels.
Chinese Tech
Xuancheng Li · Beining Wang · Haitao Li · Heng Wang · Yujia Zhou · Qingyi Pan · +4 more
Tsinghua University · Tencent
Research Digest··2 min read
Li and colleagues address a weakness in generative reward models: choosing the preferred response correctly does not guarantee that the accompanying critique is sound.
Why this paper
From Tencent and Tsinghua University
In one line
EnGRICH uses human critiques to train a MetaCritic that improves generative reward model training with process supervision and guided exploration.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§