SimpleCluster operates on visual tokens after they pass through the model's projector, which maps vision-encoder outputs into the language model's representation space.
Simple clustering preserves video understanding with far fewer visual tokens
Across four benchmarks and three video language models, a training-free clustering method matched or outperformed more elaborate compression approaches, including at 1 percent token retention.
Chinese Tech
Xiao Zhang · Wang Zeng · Sheng Jin · Wentao Liu · Chen Qian · Shichao Kan
Central South University · SenseTime Research · Tetras.AI
Research Digest··3 min read
Zhang et al.
Why this paper
From SenseTime Research and 2 others · Released code
In one line
Simple clustering of visual tokens by position-aware cross-frame features preserves structure and matches or beats complex compression methods.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§