vision-language model
A model that handles tasks combining visual and textual information, such as editing images from conversation or reasoning across images and text.
- Papers
- 25
- Released code
- 5
- First seen
- Mar 2026
- Latest
- Sept 2026
19 papers in the last two months, against 4 in the two before.
The papers
Most central to this idea first, not most recent.
- Big Techcs.CL
Agentic document QA pays off mainly with stronger vision-language models
Amazon.com · Sept 2026
- Top Universitycs.RO
Visual action rehearsal helps language models control robot manipulation
HKUST(GZ), CUHK · Sept 2026
- Big Techcs.CV
A single visual agent spans mobile, desktop, web, and tool use
Apple · Sept 2026
- Industrycs.AI
Decouple high-level VLM planning from low-level execution to speed up mobile agents
Rice University · Sept 2026
- Independentcs.RO
A semantic harness lets vision-language models control different robots
Sept 2026
- Independentcs.AI
Scene-grounded decoding keeps vision-language plans executable and visually supported
Sept 2026
- Industrycs.AIcode
Agents learn selectively from imperfect vision-language model teachers
Fondazione Bruno Kessler, University of Torino · Sept 2026
- Top Universitycs.CV
Factorized reinforcement learning improves vision-language models’ spatial reasoning
Joy Future Academy, The Hong Kong University of Science and Technology (Guangzhou) · Sept 2026
- Independentcs.CV
Coding agents combine generated imagery with editable web-based visual layouts
Sept 2026
- Chinese Techcs.CV
Conversational image editing agent learns interpretable tool use
Shanghai Innovation Institution, Huawei Technologies Ltd. · June 2026
- Independentcs.SE
Multi-image evidence can help models repair software, but unreliably
Sept 2026
- Chinese Techcs.AIcode
C3M preserves multimodal evidence across sessions within fixed memory budgets
Huzhou Normal University, Alibaba Group · Sept 2026
- Top Universitycs.RO
Saliency-driven workspace tokens give robots lightweight task memory
Massachusetts Institute of Technology, Carnegie Mellon University · Sept 2026
- Chinese Techcs.AI
Self-improving context programs help frozen models understand long videos
Tencent · Aug 2026
- Top Universitycs.CVcode
MLLMs fail to sustain goal-directed navigation in real-scale 3D city
Shanghai Jiao Tong University, National University of Singapore · Aug 2026
- Industrycs.CL
Self-evolving loop synthesizes high-quality multimodal training data
vivo AI Lab · Aug 2026
- Chinese Techcs.CV
Adaptive tool use improves video research agents’ accuracy and efficiency
Accio Team, Alibaba Group, Beijing Key Laboratory of Intelligent Information Technology · Aug 2026
- Big Techcs.AI
New framework lets you train AI agents inside the same harness systems they use at inference
Columbia University, Dartmouth College · July 2026
- Big Techcs.AI
Skill evolution improves agent performance in image generation workflows
University of Pennsylvania, Nvidia · July 2026
- Research Labcs.CR
Execution Logs Can Mislead Visual Judges in Video-Generation Agents
RIKEN · Sept 2026
- Independentcs.CVcode
Image-only pretraining helps build a strong open image generator
Sept 2026
- Chinese Techcs.CV
Streaming video memory works better when internalized as evolving latent tokens
Nanjing University of Science and Technology, Ant Group · Sept 2026
- Top Universitycs.CVcode
Multi-agent framework coordinates narrative and visual consistency for long-form video
The University of Hong Kong, South China University of Technology · June 2026
- Independentcs.CV
Code-driven agents generate images via sketching and texturing stages
May 2026
- Chinese Techcs.CV
Unified visual-generation agentic model outperforms larger closed-source models
Tencent Hunyuan, Hong Kong University of Science and Technology · Mar 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.