What they did
The authors created a robotic manipulation dataset spanning 10 task categories and 56 base tasks. Each task appears at five difficulty levels, producing 280 variants that progressively increase spatial ambiguity or procedural complexity.
RoboSPA contains 527,000 trajectories collected across multiple robot embodiments and varied scenes. Its evaluations supplement binary task success with diagnostic metrics intended to identify specific reasoning and execution failures.
Key findings
- Representative vision-language-action models performed worse as spatial and procedural complexity increased.
- Models struggled to resolve fine-grained spatial relationships in ambiguous scenes.
- Precise low-level action execution remained a significant source of failure.
- Longer tasks exposed weaknesses in memory-intensive procedural planning.
Why it matters
RoboSPA shifts evaluation from asking only whether a robot completed a task to examining where its reasoning or execution broke down. This could help researchers distinguish failures of spatial understanding, action control, and long-horizon planning when developing more reliable embodied agents.
Caveats
The abstract does not identify the evaluated models, report quantitative performance scores, or detail the diagnostic metrics, robot platforms, and scene diversity. It therefore supports the conclusion that tested systems struggled on the benchmark, but not broader claims about all vision-language-action models or real-world deployment.