Bounded visual workspaces improve multimodal agents’ accuracy and efficiency
Yang et al. introduce VLM-in-Sandbox, which stores generated crops, masks, overlays, and other image artifacts in a ledger rather than continually appending them to the model’s context. Across seven benchmarks and four vision-language models, explicit control over which images remain visible improved aggregate accuracy while reducing token use and latency.
23 Sept 2026