Theory and experiments show when splitting large images into views helps VLMs find small objects

Authors derive conditions under which local decomposition preserves recall and test them across 11 models on 797 images.

Top University
Junzhe Shi · Yuan Gan · Shida Jiang

University of California, Berkeley · Quotr AI

Research Digest··4 min read
The authors analyzed why vision-language models (VLMs) miss small objects, deriving a theory based on two quantities: S (visual tokens across an object's side) and L (content per call).

The authors formalized the VLM image interface: a model resizes an image to a token budget and cuts it into visual tokens of known size.

Why this paper

From University of California, Berkeley and Quotr AI · Released code

In one line

Small-object detection in VLMs is limited by tokens per object and content per call; overlapping, non-zoomed-out decomposition preserves recall under monotonicity assumptions.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks (7 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe