Robots recover more often when vision-language models monitor failures

Spotter lets an embodied policy control the robot continuously while a parallel vision-language model detects, diagnoses and repairs execution errors.

Chinese Tech
Long Li · Qichao Zhao · Yue Yang · Fan Xu · Zhe Wang · Alan Wee-Chung Liew · +3 more

Griffith University · IIIS, Tsinghua University · Tencent · Fudan University · Tongji University

Research Digest··3 min read
Li and colleagues test whether robot policies can recognize and correct failures during manipulation, finding that retries, textual error descriptions and best-of-N selection rarely produce reliable recovery.

The authors first separated error recovery into two abilities: generating a corrective action and judging whether that correction will work.

Why this paper

From Tencent and 4 others

In one line

Spotter improves robot task success by letting embodied policies act continuously while a parallel vision-language model detects, diagnoses, and repairs failures.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (3 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.