The authors studied Rubric-RL, where an LLM judge evaluates separate checklist criteria and combines the verdicts into a training reward.
Protocol-Level Rewards Reduce Rubric Gaming in Clinical Language Models
Requiring related criteria to pass together improved clinical appropriateness without reducing checklist coverage.
Chinese Tech
Maoqi Liu · Junwei He · Bowen Zhang · Feiran Li · Wentao Ma · Rongyi Lin · +2 more
Beijing University of Posts and Telecommunications · ByteDance
Research Digest··2 min read
Liu et al.
Why this paper
From ByteDance and Beijing University of Posts and Telecommunications
In one line
ProRubric groups rubric criteria into conjunctive protocol-level dimensions to prevent reward hacking from additive aggregation.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§