← All concepts

vision-language-action model

A robot policy model that maps visual observations and language instructions to actions, trained on demonstrations and often used for manipulation tasks.

Papers
7
Released code
1
First seen
Aug 2026
Latest
Sept 2026

7 papers in the last two months, against 0 in the two before.

The papers

Most central to this idea first, not most recent.

Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.

vision-language-action model — Zotpaper Research