The authors build on coarse-to-fine flexible tokenizers, where a video is encoded into an ordered token sequence and any prefix can be used.
Semantic token prefixes let smaller autoregressive video models match larger ones
SemanTok injects frozen DINO features into a flexible video tokenizer so early tokens carry semantics and are cheaper for an autoregressive model to predict.
AI Startup
Mikhail Dereviannykh · Vikram Voleti · Simon Donne · Mallikarjun Byrasandra Ramalinga Reddy · Shimon Vainer · Mark Boss
Stability AI · Karlsruhe Institut für Technologie
Research Digest··2 min read
Dereviannykh et al.
Why this paper
From Stability AI and Karlsruhe Institut für Technologie
In one line
SemanTok, a flexible video tokenizer, achieves high semantic alignment and video fidelity, matching or beating models 3.4 times its size.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§