Changing reward requirements teaches vision-language models to actually use images

Experiments with counterfactual scenes show that reinforcement learning improves visual grounding only when correct answers require extracting evidence from the image.

Research Lab
Haocun Ye · Xinlong Jiang · Qile Chen · Bingyu Wang · Teng Zhang · Shubai Chen · +3 more

University of the Chinese Academy of Sciences · Institute of Computing Technology, Chinese Academy of Sciences

Research Digest··3 min read
Ye and colleagues test whether reinforcement learning with verifiable rewards improves visual perception or merely teaches models to exploit nonvisual shortcuts.

The authors first examined image-based reinforcement learning on Geometry3K, comparing normal training against runs where every training image was replaced by a gray canvas.

Why this paper

From University of the Chinese Academy of Sciences and Institute of Computing Technology, Chinese Academy of Sciences

In one line

Multimodal reinforcement learning with verifiable rewards fails to learn visual skills unless tasks require visual evidence.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ✓Compute or model size stated (params 7B)
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.