Adaptive visual tokens cut multimodal inference costs while preserving accuracy

VisionWeave learns where images need fine detail and where coarse representations suffice, reducing visual token counts without relying on post-training pruning.

Chinese Tech
Yuan Feng · Qize Yang · Ruizhe Chen · Sibo Song · Haolin He · Muzhi Zhu · +6 more

University of Science and Technology of China · Alibaba Token Hub, Alibaba Group · The Chinese University of Hong Kong · Zhejiang University

Research Digest··2 min read
Feng and colleagues trained multimodal language models to allocate visual tokens at different spatial granularities according to image content.

VisionWeave augments Qwen-based multimodal models with a gated spatial pooler and a granularity router.

Why this paper

From Alibaba Token Hub, Alibaba Group and 3 others

In one line

VisionWeave learns where to use coarse or fine visual tokens, cutting average token use by 43.0% while retaining 98.9% of native performance.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe