Cheng et al.
Model predicts full 6-DoF motion at every pixel from monocular video
MoSE3 decomposes SE(3) into point tracks and rigidity embeddings, enabling end-to-end learning without category priors or per-video optimization.
Top University
Jiahuan Cheng · Zhiyi Li · Tian Xia · Ruojin Cai · Yilun Du · Qianqian Wang
Harvard University · Kempner Institute · Johns Hopkins University · MIT
Research Digest··3 min read
The authors present MoSE3, the first feed-forward model to predict dense per-pixel SE(3) (3D rotation and translation) from monocular RGB video.
Why this paper
From Harvard University and 3 others
In one line
MoSE3 predicts dense per-pixel SE(3) motion from monocular RGB video.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§