Sparse annotation overlap makes LLM judge validation decisions unreliable

The authors show that shared labeling coverage, more than the agreement metric chosen, determines whether an automated judge is approved or ranked correctly.

Big Tech
Junxuan Li · Arko Mukherjee · Soumyabrata Pal

Adobe · Adobe Research

Research Digest··2 min read
Li, Mukherjee and Pal analyze how limited overlap, the proportion of items labeled by multiple raters, affects validation of LLM judges against humans.

The authors derive how sparse overlap affects the bias and variance of several inter-rater agreement estimators, including observed agreement, AC1, Cohen's kappa and Krippendorff's alpha.

Why this paper

From Adobe Research and Adobe

In one line

Sparse overlap is the primary cause of incorrect validation decisions; 25% overlap suffices for non-borderline judges.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.