The authors first performed a fixed-context Pass@K analysis.
Self-distillation helps vision-language models reason with far fewer visual tokens
SCOPD trains models on pruned visual inputs using full-context teacher guidance, recovering part of the accuracy lost under aggressive token pruning.
Big Tech
Ahmadreza Jeddi · Enming Zhang · Jasper Gerigk · Hakki Karaimer · Mozhgan Nasr Azadani · Jiayun Luo · +9 more
University of Toronto · Vector Institute · Samsung AI Center Toronto · Stanford University · University of Waterloo
Research Digest··2 min read
Jeddi et al.
Why this paper
From Samsung AI Center Toronto and 9 others
In one line
Visual token pruning degrades VLM performance primarily due to unreliable use of remaining visual evidence, not irreversible information loss.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§