, Qwen-VL, LLaVA, InternVL, NVILA).
Visual attention sinks in VLM decoders are layer-dependent and can be steered
SAGE identifies prompt-invariant sinks in early and late layers and improves grounding by redirecting attention to query-relevant regions.
Academic
Jeonghyo Song · YoungJoon Yoo
Chung-Ang University
Research Digest··3 min read
The authors analyze decoder attention in vision-language models and discover that early and late layers consistently attend to the same image regions regardless of the prompt, while middle layers shift attention meaningfully.
Why this paper
From Chung-Ang University
In one line
SAGE improves visual grounding and reduces hallucinations by steering decoder attention away from prompt-invariant sinks in early and late layers.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§