The authors built S4VY on a feed-forward visual-geometry backbone that extracts jointly contextualized geometric features directly from a set of RGB observations.
S4VY segments persistent objects across unordered views of dynamic scenes
The model uses shared spatial and temporal geometry to produce class-agnostic object masks, accept prompts and maintain identities across RGB observations.
Big Tech
Jingdong Zhang · Xin Li · Jan Kautz · Wenping Wang · Chris Choy
Texas A&M University, College Station · NVIDIA
Research Digest··2 min read
Zhang et al.
Why this paper
From NVIDIA and Texas A&M University, College Station
In one line
S4VY segments any object in 4D scenes from raw images using feed-forward visual geometry, supporting prompts and language grounding.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§