VisionWeave augments Qwen-based multimodal models with a gated spatial pooler and a granularity router.
Adaptive visual tokens cut multimodal inference costs while preserving accuracy
VisionWeave learns where images need fine detail and where coarse representations suffice, reducing visual token counts without relying on post-training pruning.
Chinese Tech
Yuan Feng · Qize Yang · Ruizhe Chen · Sibo Song · Haolin He · Muzhi Zhu · +6 more
University of Science and Technology of China · Alibaba Token Hub, Alibaba Group · The Chinese University of Hong Kong · Zhejiang University
Research Digest··2 min read
Feng and colleagues trained multimodal language models to allocate visual tokens at different spatial granularities according to image content.
Why this paper
From Alibaba Token Hub, Alibaba Group and 3 others
In one line
VisionWeave learns where to use coarse or fine visual tokens, cutting average token use by 43.0% while retaining 98.9% of native performance.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§