The authors first separated error recovery into two abilities: generating a corrective action and judging whether that correction will work.
Robots recover more often when vision-language models monitor failures
Spotter lets an embodied policy control the robot continuously while a parallel vision-language model detects, diagnoses and repairs execution errors.
Chinese Tech
Long Li · Qichao Zhao · Yue Yang · Fan Xu · Zhe Wang · Alan Wee-Chung Liew · +3 more
Griffith University · IIIS, Tsinghua University · Tencent · Fudan University · Tongji University
Research Digest··3 min read
Li and colleagues test whether robot policies can recognize and correct failures during manipulation, finding that retries, textual error descriptions and best-of-N selection rarely produce reliable recovery.
Why this paper
From Tencent and 4 others
In one line
Spotter improves robot task success by letting embodied policies act continuously while a parallel vision-language model detects, diagnoses, and repairs failures.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§