vision-language-action model
A robot policy model that maps visual observations and language instructions to actions, trained on demonstrations and often used for manipulation tasks.
- Papers
- 7
- Released code
- 1
- First seen
- Aug 2026
- Latest
- Sept 2026
7 papers in the last two months, against 0 in the two before.
The papers
Most central to this idea first, not most recent.
- Top Universitycs.RO
Synthetic demonstrations let robot policies escape sparse-reward failures
Fujitsu Limited, The Institute of Statistical Mathematics · Sept 2026
- Industrycs.RO
Adaptive action chunking improves robot control by varying horizon based on prediction reliability.
Tongji University, The Hong Kong Polytechnic University · Sept 2026
- Independentcs.RO
Current robot models falter as spatial and procedural complexity rises
Sept 2026
- Chinese Techcs.CV
Language enables precise character and camera control in video worlds
Tencent, National University of Singapore · Sept 2026
- Top Universitycs.RO
Agent-side memory steers stateless robots through long manipulation tasks
KU Leuven, Meituan Inc. · Sept 2026
- Top Universitycs.RO
Handheld demonstrations improve robot policies without repeated robot execution
Xi’an Jiaotong University, National Key Laboratory for Multimedia Information Processing · Sept 2026
- Top Universitycs.ROcode
One video world model transfers physical dynamics across robot bodies
Princeton University · Aug 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.