The authors built Mosaic, a visual harness through which a multimodal large language model can construct intermediate views using ten composable operations.
Visual image rearrangement helps models when exact evidence matters most
A ten-operation visual harness improved fine-grained multi-image reasoning selectively, and an 8B model learned to compose its tools through reinforcement learning.
Industry
Gengyuan Zhang · Xiao Han · Xinyu Xie · Tong Liu · Volker Tresp
LMU Munich · MCML
Research Digest··2 min read
Zhang et al.
Why this paper
From LMU Munich and MCML
In one line
Visual re-representation delivers the strongest benefits when multi-image reasoning depends on precise comparisons, transformations, orientations, or spatial evidence.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§