The authors constructed ScribbleEdit with an automated pipeline that generates scribble-image pairs covering distinct editing operations, such as adding, removing or altering objects and attributes.
Benchmark shows image editing models fail to parse scribble-only intent
Across seven vision-language editing models, the authors find that scribble intention understanding, not generation ability, is the main failure point.
Chinese Tech
Jie Ren · Hao Kang · Kai Guo · Yiding Yang · Bo Liu · Liming Jiang · +6 more
MIT · ByteDance · Michigan State University
Research Digest··2 min read
The authors introduce ScribbleEdit, a benchmark for image editing in which the user's only input is a scribble drawn on the image.
Why this paper
From ByteDance and 2 others
In one line
Existing VLM/LLM-based image editing models fail to interpret scribble-only editing intents; a soft-token baseline improves understanding.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§