The authors introduce VR-JEPA, which extends the Video Joint-Embedding Predictive Architecture (V-JEPA) for generation-based video reasoning.
Latent trajectory guidance improves video generation reasoning
VR-JEPA uses contrastive learning on V-JEPA features to align predicted state sequences with task logic.
Chinese Tech
Zehua Ma · Kun Xiang · Yunshuang Nie · Quanlin Chen · Haoyuan Li · Xiuwei Chen · +6 more
Sun Yat-sen University · Shenzhen Loop Area Institute · Tsinghua Shenzhen International Graduate School · Tencent · Mohamed bin Zayed University of Artificial Intelligence
Research Digest··2 min read
The authors propose VR-JEPA, a framework that predicts task-relevant latent state trajectories using V-JEPA features and uses them to guide a video diffusion model.
Why this paper
From Tencent and 5 others
In one line
VR-JEPA uses latent trajectory prediction to guide video generation for improved visual reasoning.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§