The authors built a multi-agent robot-control harness around general-purpose vision-language models.
Visual action rehearsal helps language models control robot manipulation
A multi-agent harness lets vision-language models preview, revise, and visually correct robot actions before and during execution.
Top University
Yehang Zhang · Haojian Huang · Yifan Chang · Jianchong Su · Bohan Zhou · Yingjie Xu · +10 more
HKUST(GZ) · CUHK · Knowin AI
Research Digest··2 min read
Zhang et al.
Why this paper
From HKUST(GZ) and 2 others · Part of World Model Planning for Agents, now 19 papers
In one line
World Action Agent achieves 75.6% success on LIBERO-Pro by having VLMs rehearse actions in a visual workspace.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§