5-VL-3B and 7B Instruct models on EgoIntention and RefEgo-Int, a new benchmark curated from RefEgo.
Probability rewards keep visual grounding reasoning aligned with predicted objects
Rita rewards reasoning that raises the likelihood of the reference answer and couples that signal to bounding-box accuracy.
Big Tech
Pengzhan Sun · Shiu-hong Kao · Shijie Li · Yongyi Su · Junbin Xiao · Arjun Reddy Akula · +1 more
National University of Singapore · A*STAR Institute of Advanced Intelligence and Computing, Singapore · South China University of Technology · University of Science and Technology of China · Google DeepMind
Research Digest··2 min read
Sun et al.
Why this paper
From Google DeepMind and 4 others
In one line
Rita reduces thinking drift in visual intention grounding by using conditional probability rewards to enforce thinking-answer consistency.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§