Self-Saliency divides a vision-language model’s generated reasoning into steps, identifies object references in each observation step, and uses an external grounding model to locate those objects in the image.
Self-grounded attention improves visual reasoning across diverse benchmarks
Training a vision-language model to attend to regions named in its own reasoning outperformed methods that select regions from the image and question alone.
Big Tech
Uri Berger · Gal Chechik · Gal Dalal
NVIDIA Research · The Hebrew University of Jerusalem · University of Melbourne · Bar-Ilan University
Research Digest··3 min read
Berger, Chechik and Dalal introduce Self-Saliency, a post-training method that grounds each step of a model’s reasoning in corresponding image regions and rewards alignment with its visual attention.
Why this paper
From NVIDIA Research and 3 others
In one line
Training VLMs to attend to objects they mention in reasoning improves visual reasoning accuracy.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§