The authors curated SubjectiveSet, a dataset of 50,013 response pairs drawn from 17 public sources, each containing two responses to the same query.
Framework reveals that LLM judges share perception but differ in prioritization
By separating how judges compare responses on attributes from how they weigh those attributes, researchers can steer subjective preferences without retraining.
Top University
Qi Cao · Kangning Liu · Xuan Kan · Shunwen Tan · Yang Pei · Dake Chen · +7 more
Meta · University of California San Diego · The University of Hong Kong · The Hong Kong University of Science and Technology
Research Digest··3 min read
The authors introduce JudgeProfile, a framework that decomposes LLM evaluation into perception (how a judge compares responses on specific qualities like clarity and correctness) and prioritization (how much each quality influences the final choice).
Why this paper
From The University of Hong Kong and 3 others
In one line
Reweighting attribute judgments steers LLM judges' subjectivity, improving target preference agreement without model updates.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§