The authors first examined image-based reinforcement learning on Geometry3K, comparing normal training against runs where every training image was replaced by a gray canvas.
Changing reward requirements teaches vision-language models to actually use images
Experiments with counterfactual scenes show that reinforcement learning improves visual grounding only when correct answers require extracting evidence from the image.
Research Lab
Haocun Ye · Xinlong Jiang · Qile Chen · Bingyu Wang · Teng Zhang · Shubai Chen · +3 more
University of the Chinese Academy of Sciences · Institute of Computing Technology, Chinese Academy of Sciences
Research Digest··3 min read
Ye and colleagues test whether reinforcement learning with verifiable rewards improves visual perception or merely teaches models to exploit nonvisual shortcuts.
Why this paper
From University of the Chinese Academy of Sciences and Institute of Computing Technology, Chinese Academy of Sciences
In one line
Multimodal reinforcement learning with verifiable rewards fails to learn visual skills unless tasks require visual evidence.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 7B)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§