Vision models weigh surface objects against abstract relations in separate circuits

Controlled tests suggest that early representations favor object similarity, while later processing can encode relational structure.

Big Tech
Gouki Minegishi · Hiroki Furuta · Takeshi Kojima · Yusuke Iwasawa · Yutaka Matsuo

The University of Tokyo · Google DeepMind

Research Digest··3 min read
Minegishi and colleagues tested whether vision language models generalize from visible objects or from the abstract relations among them.

The authors adapted Relational Match-to-Sample, a paradigm from developmental psychology.

Why this paper

From Google DeepMind and The University of Tokyo

In one line

VLMs implement abstract reasoning via two competing circuits: an early object-feature circuit and a late relation circuit.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.