The authors assembled 92 tasks from 36 open-source organizations, split into 48 public tasks and 44 private holdouts.
Multi-image evidence can help models repair software, but unreliably
SWE-PolyVision tests whether coding models can combine clues across multiple visuals and text to produce repository-level repairs that pass isolated verification.
Independent
Jiajun Wu · Leixin Sun · Zihan Tan · Yitao Liu · Shuo Li · Jiaru Qian · +6 more
Research Digest··2 min read
Thread:Live Software Adaptation
The authors introduce an executable benchmark containing 92 software-repair tasks, each with at least two visual inputs and a fixed pre-repair repository.
Why this paper
Independent · Part of Live Software Adaptation, now 4 papers
In one line
SWE-PolyVision benchmarks cross-image reasoning for repository-level repair and finds visual access effects are model- and task-dependent.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§