VETO applies compression between a model’s visual encoder and language model, without requiring architectural changes or custom CUDA kernels.
Compressing Video Tokens Across Space and Time Speeds Vision-Language Models
VETO first merges redundant tokens within frames, then merges similar frames, reducing inference time while largely preserving model accuracy.
Big Tech
Gueter Josmy Faure · Hao Ping Wang · Min-Hung Chen · Winston H. Hsu
National Taiwan University · NVIDIA
Research Digest··2 min read
Faure et al.
Why this paper
From NVIDIA and National Taiwan University
In one line
VETO compresses video tokens with hierarchical intra-frame and inter-frame merging, achieving up to 45% faster inference in VLMs while preserving accuracy.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (4 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§