The authors identified that standard RLVR with GRPO uses a single chain-of-thought and final-answer reward, which cannot isolate errors in visual perception versus language reasoning.
Separating visual and language reasoning phases improves vision-language model accuracy
SPLIT-RL trains visual perception and language reasoning in disjoint stages, using claim-level rewards to refine each capability individually.
Big Tech
Raja Kumar · Rajat Koner · Ritwick Chaudhry · Zhuowei Li · Nishant Sankaran · Yifan Xing
University of Southern California · Amazon AGI
Research Digest··2 min read
The authors propose SPLIT-RL, a staged reinforcement learning method that trains a vision-language model's visual reasoning (VR) and language reasoning (LR) capabilities in separate phases.
Why this paper
From Amazon AGI and University of Southern California
In one line
SPLIT-RL separates visual and language reasoning training stages with claim-level advantages, improving accuracy over GRPO by 1.4-6.1 points.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§