← All concepts

vision-language model

A model that handles tasks combining visual and textual information, such as editing images from conversation or reasoning across images and text.

Papers
25
Released code
5
First seen
Mar 2026
Latest
Sept 2026

19 papers in the last two months, against 4 in the two before.

The papers

Most central to this idea first, not most recent.

Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.