Framework reveals that LLM judges share perception but differ in prioritization

By separating how judges compare responses on attributes from how they weigh those attributes, researchers can steer subjective preferences without retraining.

Top University
Qi Cao · Kangning Liu · Xuan Kan · Shunwen Tan · Yang Pei · Dake Chen · +7 more

Meta · University of California San Diego · The University of Hong Kong · The Hong Kong University of Science and Technology

Research Digest··3 min read
The authors introduce JudgeProfile, a framework that decomposes LLM evaluation into perception (how a judge compares responses on specific qualities like clarity and correctness) and prioritization (how much each quality influences the final choice).

The authors curated SubjectiveSet, a dataset of 50,013 response pairs drawn from 17 public sources, each containing two responses to the same query.

Why this paper

From The University of Hong Kong and 3 others

In one line

Reweighting attribute judgments steers LLM judges' subjectivity, improving target preference agreement without model updates.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.