Hidden prompt revisions inject cultural stereotypes into generated images
The authors introduce WORLDVIEW, a benchmark of 8,960 prompts spanning 15 languages and 31 language–context pairings, to examine how text-to-image systems rewrite requests before generation. Across DALL-E-3, Imagen-4, and GPT-Image-1.5, revisions marked non-Western and non-Anglophone settings more heavily than the United States and repeatedly reduced them to narrow, stereotypical vocabularies.
13 Sept 2026