The authors adapted Relational Match-to-Sample, a paradigm from developmental psychology.
Vision models weigh surface objects against abstract relations in separate circuits
Controlled tests suggest that early representations favor object similarity, while later processing can encode relational structure.
Big Tech
Gouki Minegishi · Hiroki Furuta · Takeshi Kojima · Yusuke Iwasawa · Yutaka Matsuo
The University of Tokyo · Google DeepMind
Research Digest··3 min read
Minegishi and colleagues tested whether vision language models generalize from visible objects or from the abstract relations among them.
Why this paper
From Google DeepMind and The University of Tokyo
In one line
VLMs implement abstract reasoning via two competing circuits: an early object-feature circuit and a late relation circuit.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§