LLM judges often encode better preferences than their verdicts reveal

Across open-weight evaluators, internal activations recovered human-aligned judgments that surface-biased final answers frequently obscured.

Big Tech
Sourabrata Mukherjee · Sunayana Sitaram

Microsoft Research

Research Digest··3 min read
Mukherjee and Sitaram test whether an incorrect LLM judgment reflects missing knowledge or a failure to express information already represented inside the model.

The authors compared judges’ final verdicts with small supervised probes trained on their frozen internal activations, without updating model weights.

Why this paper

From Microsoft Research

In one line

LLM judges' internal representations often contain correct evaluations even when their final verdict is wrong.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.