The authors start from the observation that humans do not maintain pixel-perfect 3D geometry, but rather roughly identify objects across views and assemble a coarse layout.
Multimodal LLMs reason better in 3D when trained to imagine a coarse scene first
Imagine3D-LLM appends learnable summary tokens to image tokens, decodes them into compact 3D Gaussians, and outperforms prior 3D-aware MLLMs on seven benchmarks.
Big Tech
Jaewoo Jung · Hyeonseo Yu · Honggyu An · Jisang Han · Mungyeom Kim · Minkyeong Jeon · +7 more
KAIST AI · ETH Zürich · Google · TUM · ETH AI Center
Research Digest··3 min read
The authors introduce Imagine3D-LLM, a multimodal large language model that learns to assemble a compact 3D scene representation before answering spatial questions.
Why this paper
From Google and 4 others
In one line
A compact, reconstruction-supervised 3D scene representation improves multi-view MLLM spatial reasoning more than injecting fine-grained pixel geometry.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§