The authors formulate 3D spatial relation segmentation in a feed-forward, pose-free multi-view setting.
Feed-forward model segments objects by spatial relations across views
RelationVGGT merges semantic and geometric features to answer relational queries without per-scene optimization or known camera poses.
Big Tech
Minsu Kim · Jaesung Choe · Jiwoo Lee · Yu-Chiang Frank Wang · Seon Joo Kim
Yonsei University · NVIDIA
Research Digest··2 min read
The authors propose RelationVGGT, a feed-forward framework that takes multi-view images, a subject mask and a relational text query, and segments the target object across all views.
Why this paper
From NVIDIA and Yonsei University
In one line
RelationVGGT performs pose-free, feed-forward multi-view segmentation of objects specified by their spatial relation to a visually marked subject.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (4 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§