Self-distillation helps vision-language models reason with far fewer visual tokens

SCOPD trains models on pruned visual inputs using full-context teacher guidance, recovering part of the accuracy lost under aggressive token pruning.

Big Tech
Ahmadreza Jeddi · Enming Zhang · Jasper Gerigk · Hakki Karaimer · Mozhgan Nasr Azadani · Jiayun Luo · +9 more

University of Toronto · Vector Institute · Samsung AI Center Toronto · Stanford University · University of Waterloo

Research Digest··2 min read
Jeddi et al.

The authors first performed a fixed-context Pass@K analysis.

Why this paper

From Samsung AI Center Toronto and 9 others

In one line

Visual token pruning degrades VLM performance primarily due to unreliable use of remaining visual evidence, not irreversible information loss.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.