The authors first examined how visual representations change across a Vision Transformer.
Pruning visual tokens and decoder sublayers cuts multimodal model computation
SPIDER combines information from multiple vision-encoder layers with selective attention and feed-forward skipping, without additional training.
Chinese Tech
Tianxiang Chen · Zhentao Tan · Zi Ye · Yue Wu · Xiaobing Tu · Jinkui Ren · +6 more
Alibaba Cloud · University of Science and Technology of China · Maynooth University · Fudan University · Alibaba Group
Research Digest··2 min read
Chen and colleagues target two sources of multimodal model overhead: redundant visual tokens and unnecessary computation inside the language-model decoder.
Why this paper
From Alibaba Cloud and 4 others
In one line
SPIDER jointly prunes visual tokens and skips decoder sublayers without training, cutting LLaVA-NeXT-7B FLOPs by 79% while retaining 96% of baseline performance.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§