IntrinSync represents an RGB image and its intrinsic maps as a shared sequence of tokens in a Multimodal Diffusion Transformer.
Joint inverse and forward rendering produces more coherent intrinsic scene maps
IntrinSync predicts five interdependent scene properties together, then uses reciprocal rendering to make them collectively explain the input image.
Chinese Tech
Zheng Gu · Rui Huang · Xilu Zhang · Jingbo Zhang · Min Lu · Zhida Sun · +3 more
Shenzhen University · Robotics X Lab, Tencent · Hebrew University of Jerusalem · Tel Aviv University
Research Digest··3 min read
Gu and colleagues present a diffusion-based system that jointly separates a single RGB image into albedo, shading, surface normals, roughness and metallic maps, while also learning to render those maps back into an image.
Why this paper
From Robotics X Lab, Tencent and 3 others
In one line
Jointly modeling intrinsic channels and closing the inverse-forward rendering loop with a dual cycle-consistent objective improves intrinsic decomposition and forward rendering coherence.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (3 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§