The authors assembled a curated dataset of icons representing abstract concepts through culturally conventional visual proxies.
Vision transformers link concrete visual cues to abstract concepts
Circuit tracing in CLIP and DINO encoders found that interpretable object-like features mediate abstract predictions, while text-bearing icons follow a separate route.
Big Tech
Jing Ding · Ziqiao Ma · Jiayuan Mao · Joyce Chai · Freda Shi
University of Michigan · Vector Institute · University of Pennsylvania · Amazon FAR · University of Waterloo
Research Digest··2 min read
Ding et al.
Why this paper
From Amazon FAR and 4 others
In one line
Vision Transformers ground abstract concepts through metonymic circuits, using concrete visual anchors like fire to bridge image features to abstract meaning.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§