A single corrupted observation can derail vision-language robot policies

Tests across five models show that brief sensor corruption is especially damaging at critical moments, while adaptive action execution can reduce the harm.

Research Lab
Shojiro Yamabe · Jun Sakuma

Science Tokyo · RIKEN

Research Digest··2 min read
Yamabe and Sakuma test vision-language-action models by corrupting exactly one observation per episode with eight simulated sensor failures.

The authors evaluated five vision-language-action models using eight one-step perturbations, including visual corruption, packet loss, abnormal state values and missing data.

Why this paper

From RIKEN and Science Tokyo

In one line

A single corrupted observation can derail VLA robot tasks, while consistency-triggered shortening of action execution chunks improves robustness without sacrificing clean performance.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.