The authors compared judges’ final verdicts with small supervised probes trained on their frozen internal activations, without updating model weights.
LLM judges often encode better preferences than their verdicts reveal
Across open-weight evaluators, internal activations recovered human-aligned judgments that surface-biased final answers frequently obscured.
Big Tech
Sourabrata Mukherjee · Sunayana Sitaram
Microsoft Research
Research Digest··3 min read
Mukherjee and Sitaram test whether an incorrect LLM judgment reflects missing knowledge or a failure to express information already represented inside the model.
Why this paper
From Microsoft Research
In one line
LLM judges' internal representations often contain correct evaluations even when their final verdict is wrong.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§