The authors created OR-Bench, comprising five two-view tasks and three multi-view tasks.
Vision-language models detect object rotation but misjudge its magnitude
OR-Bench finds that models retain usable coarse rotation information, while a lightweight decoder helps them access it.
Independent
Zhaochen Wang · Yujun Cai · Huangbo Zou · Hower Yang · Naipeng Dong · Miao Xu · +1 more
Research Digest··2 min read
Wang and colleagues evaluated 12 vision-language models on eight tasks requiring them to detect, estimate and reason about object rotations across views.
Why this paper
Independent
In one line
VLMs detect object rotation but poorly estimate its magnitude; reinjecting decoded rotation cues from frozen representations improves accuracy by 7.9 to 12.6 points.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§