Simple clustering preserves video understanding with far fewer visual tokens

Across four benchmarks and three video language models, a training-free clustering method matched or outperformed more elaborate compression approaches, including at 1 percent token retention.

Chinese Tech
Xiao Zhang · Wang Zeng · Sheng Jin · Wentao Liu · Chen Qian · Shichao Kan

Central South University · SenseTime Research · Tetras.AI

Research Digest··3 min read
Zhang et al.

SimpleCluster operates on visual tokens after they pass through the model's projector, which maps vision-encoder outputs into the language model's representation space.

Why this paper

From SenseTime Research and 2 others · Released code

In one line

Simple clustering of visual tokens by position-aware cross-frame features preserves structure and matches or beats complex compression methods.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.