5-397B) as annotators for two MT evaluation schemes: Multidimensional Quality Metrics (MQM), which assigns fine-grained error categories and severities, and Error Span Annotation (ESA), which uses severity only.
LLMs match human annotators on some MT quality tasks, but both remain unreliable
Across 70 language pairs and WMT23/WMT25 data, LLM-human agreement sometimes exceeds human-human agreement, yet varies widely by language pair, domain and annotation scheme.
Big Tech
Hala Almaghout · Christian Federmann · Qin Gao
Apple
Research Digest··2 min read
The authors evaluated four LLMs as annotators for two machine translation quality schemes, MQM and ESA, comparing their agreement with human annotators across a 70-language-pair test set and WMT23/WMT25 data.
Why this paper
From Apple
In one line
LLMs match human-human agreement on some MT quality annotation tasks, but both are too unreliable to replace human evaluation.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§