The authors evaluated five vision-language-action models using eight one-step perturbations, including visual corruption, packet loss, abnormal state values and missing data.
A single corrupted observation can derail vision-language robot policies
Tests across five models show that brief sensor corruption is especially damaging at critical moments, while adaptive action execution can reduce the harm.
Research Lab
Shojiro Yamabe · Jun Sakuma
Science Tokyo · RIKEN
Research Digest··2 min read
Yamabe and Sakuma test vision-language-action models by corrupting exactly one observation per episode with eight simulated sensor failures.
Why this paper
From RIKEN and Science Tokyo
In one line
A single corrupted observation can derail VLA robot tasks, while consistency-triggered shortening of action execution chunks improves robustness without sacrificing clean performance.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§