7M clips from OpenVid and Moments-in-Time v2 for 8 epochs.
Decoupled temporal axis yields motion-prioritized video representations efficiently
TT-VidT combines per-frame spatial features with a compact temporal path and Diff Compression to outperform prior methods on motion-sensitive tasks while using fewer FLOPs.
Big Tech
Shih-Ying Yeh · Daniel Z. Kaplan · Xuehai Wang · Fu-En Yang · Min-Hung Chen · Shang-Hong Lai
National Tsing Hua University · Comfy Org Research · realiz.ai · Karolinska Institutet · Stockholm University
Research Digest··2 min read
The authors conduct a controlled architecture-objective study and propose TT-VidT, which uses a DINOv3-initialized ViT-B/16 per-frame spatial path and a compact Temporal Transfer Layer, trained with Diff Compression to reconstruct frames from a first-frame anchor and motion tokens.
Why this paper
From NVIDIA and 5 others
In one line
TT-VidT produces video representations that prioritize motion by using a compact temporal bottleneck with diffusion-based reconstruction.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (4 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§