Decoupled temporal axis yields motion-prioritized video representations efficiently

TT-VidT combines per-frame spatial features with a compact temporal path and Diff Compression to outperform prior methods on motion-sensitive tasks while using fewer FLOPs.

Big Tech
Shih-Ying Yeh · Daniel Z. Kaplan · Xuehai Wang · Fu-En Yang · Min-Hung Chen · Shang-Hong Lai

National Tsing Hua University · Comfy Org Research · realiz.ai · Karolinska Institutet · Stockholm University

Research Digest··2 min read
The authors conduct a controlled architecture-objective study and propose TT-VidT, which uses a DINOv3-initialized ViT-B/16 per-frame spatial path and a compact Temporal Transfer Layer, trained with Diff Compression to reconstruct frames from a first-frame anchor and motion tokens.

7M clips from OpenVid and Moments-in-Time v2 for 8 epochs.

Why this paper

From NVIDIA and 5 others

In one line

TT-VidT produces video representations that prioritize motion by using a compact temporal bottleneck with diffusion-based reconstruction.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks (4 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.