The authors formalized the VLM image interface: a model resizes an image to a token budget and cuts it into visual tokens of known size.
Theory and experiments show when splitting large images into views helps VLMs find small objects
Authors derive conditions under which local decomposition preserves recall and test them across 11 models on 797 images.
Top University
Junzhe Shi · Yuan Gan · Shida Jiang
University of California, Berkeley · Quotr AI
Research Digest··4 min read
The authors analyzed why vision-language models (VLMs) miss small objects, deriving a theory based on two quantities: S (visual tokens across an object's side) and L (content per call).
Why this paper
From University of California, Berkeley and Quotr AI · Released code
In one line
Small-object detection in VLMs is limited by tokens per object and content per call; overlapping, non-zoomed-out decomposition preserves recall under monotonicity assumptions.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (7 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§