VICO jointly trains two vision-language models.
Co-evolving visual tasks keeps model training near its learning frontier
VICO jointly trains a vision-language model and a visual task rewriter that adjusts difficulty using the model’s observed success rate.
Big Tech
Meng Lu · Ligeng Zhu · Olivia Xiao · Yuchen Zhuang · Zihan Wang · Kuncheng Wu · +5 more
Virginia Tech · Cisco · NVIDIA · UChicago · Georgia Tech
Research Digest··3 min read
Lu and colleagues address a weakness of reinforcement learning with verifiable rewards: fixed training tasks eventually become either trivial or effectively impossible, leaving little useful feedback.
Why this paper
From NVIDIA and 8 others
In one line
Co-evolving visual environments with vision-language models via rollout pass rates improves out-of-domain reasoning by up to 5%.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§