Compressing Video Tokens Across Space and Time Speeds Vision-Language Models

VETO first merges redundant tokens within frames, then merges similar frames, reducing inference time while largely preserving model accuracy.

Big Tech
Gueter Josmy Faure · Hao Ping Wang · Min-Hung Chen · Winston H. Hsu

National Taiwan University · NVIDIA

Research Digest··2 min read
Faure et al.

VETO applies compression between a model’s visual encoder and language model, without requiring architectural changes or custom CUDA kernels.

Why this paper

From NVIDIA and National Taiwan University

In one line

VETO compresses video tokens with hierarchical intra-frame and inter-frame merging, achieving up to 45% faster inference in VLMs while preserving accuracy.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks (4 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.