The authors propose VIEScore2, a unified evaluator that represents images as an N×N grid and predicts quality scores and defect locations (visual artifacts or semantic misalignments) in a single pass of a VLM.
A unified evaluator scores images and pinpoints defects with explanations
The model jointly predicts quality scores and defect cells on a grid, outperforming general-purpose VLMs on score correlation and matching specialized methods on localization.
Big Tech
Xianda Du · Max Ku · Weiming Ren · Zhi Rui Tam · Chunlin Ren · Ping Nie · +2 more
University of Waterloo · NVIDIA · National Taiwan University · Nanyang Technological University
Research Digest··2 min read
The authors trained a vision-language model (VLM) on 38,000 examples with score and localization supervision, then applied group relative policy optimization (GRPO) to improve defect localization.
Why this paper
From NVIDIA and 3 others
In one line
VIEScore2 jointly predicts image quality scores and spatially grounded defect locations, outperforming general-purpose VLMs on both tasks.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§