The authors designed ViMoD, comprising two components.
Adaptive visual memory cuts multimodal reasoning costs without losing accuracy
ViMoD maintains a compact visual context and selectively recalls fine-grained evidence as reasoning needs evolve, outperforming one-shot compression baselines at a 20% token budget.
Chinese Tech
Yicheng Xue · Han Wu · Jufeng Yang · Minjing Dong · Xinghao Chen · Hanting Chen · +1 more
City University of Hong Kong · Zhejiang University · Peking University · Nankai University · Huawei Technologies
Research Digest··2 min read
The authors introduce ViMoD, a framework that couples learned visual memory with adaptive access to fine-grained tokens, targeting a low average active visual token budget during decoding.
Why this paper
From Huawei Technologies and 4 others
In one line
Adaptive selection of fine-grained visual tokens during decoding improves multimodal reasoning accuracy and efficiency compared with one-shot compression.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params Qwen3-VL-4B)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§