Pruning visual tokens and decoder sublayers cuts multimodal model computation

SPIDER combines information from multiple vision-encoder layers with selective attention and feed-forward skipping, without additional training.

Chinese Tech
Tianxiang Chen · Zhentao Tan · Zi Ye · Yue Wu · Xiaobing Tu · Jinkui Ren · +6 more

Alibaba Cloud · University of Science and Technology of China · Maynooth University · Fudan University · Alibaba Group

Research Digest··2 min read
Chen and colleagues target two sources of multimodal model overhead: redundant visual tokens and unnecessary computation inside the language-model decoder.

The authors first examined how visual representations change across a Vision Transformer.

Why this paper

From Alibaba Cloud and 4 others

In one line

SPIDER jointly prunes visual tokens and skips decoder sublayers without training, cutting LLaVA-NeXT-7B FLOPs by 79% while retaining 96% of baseline performance.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.