Self-grounded attention improves visual reasoning across diverse benchmarks

Training a vision-language model to attend to regions named in its own reasoning outperformed methods that select regions from the image and question alone.

Big Tech
Uri Berger · Gal Chechik · Gal Dalal

NVIDIA Research · The Hebrew University of Jerusalem · University of Melbourne · Bar-Ilan University

Research Digest··3 min read
Berger, Chechik and Dalal introduce Self-Saliency, a post-training method that grounds each step of a model’s reasoning in corresponding image regions and rewards alignment with its visual attention.

Self-Saliency divides a vision-language model’s generated reasoning into steps, identifies object references in each observation step, and uses an external grounding model to locate those objects in the image.

Why this paper

From NVIDIA Research and 3 others

In one line

Training VLMs to attend to objects they mention in reasoning improves visual reasoning accuracy.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.