Sparse Human Label Overlap Undermines Reliable LLM Judge Validation

The authors show that shared annotation coverage, more than the chosen agreement metric, determines whether validation produces correct deployment and ranking decisions.

Big Tech
Junxuan Li · Arko Mukherjee · Soumyabrata Pal

Adobe · Adobe Research

Research Digest··3 min read
Li, Mukherjee and Pal study how reliably teams can validate an LLM judge when only a small fraction of items receive labels from both humans and the model.

The authors formalized judge validation as a decision problem based on estimated human-model agreement.

Why this paper

From Adobe Research and Adobe

In one line

Sparse human-annotator overlap drives wrong LLM-judge deployment decisions: at 5% pairwise overlap error is 25%, and 0.25 overlap suffices for non-borderline judges.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks (3 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.