unified multimodal understanding and generation
A single end-to-end model processes and generates multiple data types (e.g., images, text) with equal competence.
- Papers
- 3
- Released code
- 2
- First seen
- Mar 2026
- Latest
- Sept 2026
The papers
Most central to this idea first, not most recent.
- Independentcs.CVcode
Image-only pretraining helps build a strong open image generator
Sept 2026
- Independentcs.CVcode
Open-source image model rivals closed systems on minimal budget
July 2026
- Chinese Techcs.CV
Unified visual-generation agentic model outperforms larger closed-source models
Tencent Hunyuan, Hong Kong University of Science and Technology · Mar 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.