← All concepts

unified multimodal understanding and generation

A single end-to-end model processes and generates multiple data types (e.g., images, text) with equal competence.

Papers
3
Released code
2
First seen
Mar 2026
Latest
Sept 2026

The papers

Most central to this idea first, not most recent.

Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.