The authors derive how sparse overlap affects the bias and variance of several inter-rater agreement estimators, including observed agreement, AC1, Cohen's kappa and Krippendorff's alpha.
Sparse annotation overlap makes LLM judge validation decisions unreliable
The authors show that shared labeling coverage, more than the agreement metric chosen, determines whether an automated judge is approved or ranked correctly.
Big Tech
Junxuan Li · Arko Mukherjee · Soumyabrata Pal
Adobe · Adobe Research
Research Digest··2 min read
Li, Mukherjee and Pal analyze how limited overlap, the proportion of items labeled by multiple raters, affects validation of LLM judges against humans.
Why this paper
From Adobe Research and Adobe
In one line
Sparse overlap is the primary cause of incorrect validation decisions; 25% overlap suffices for non-borderline judges.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§