Visual evidence works better when tied to the candidate it supports

CROSS-Bench tests whether evidence effects disappear when invalidated and move appropriately when rebound to a different answer.

Research Lab
Huiyao Zhang · Jin Bai · Zilong Su · Rui Guo · Chaofan Qin · Jinze Lv · +3 more

Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences · University of Chinese Academy of Sciences · University of Science and Technology of China

Research Digest··3 min read
Zhang and colleagues study whether vision-language models direct visual evidence toward the answer candidate that the evidence actually supports, rather than merely producing the correct final answer.

The authors introduce candidate-bound visual contribution, a framework separating three properties: utility, whether valid evidence helps; specificity, whether its added effect disappears when its supporting relation is invalidated; and ownership, whether that effect moves when the evidence is validly rebound to another candidate.

Why this paper

From Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences and 2 others

In one line

Visual evidence can improve task accuracy while its effect supports the wrong candidate; RIVET raises effect transfer and accuracy.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.