Benchmark shows image editing models fail to parse scribble-only intent

Across seven vision-language editing models, the authors find that scribble intention understanding, not generation ability, is the main failure point.

Chinese Tech
Jie Ren · Hao Kang · Kai Guo · Yiding Yang · Bo Liu · Liming Jiang · +6 more

MIT · ByteDance · Michigan State University

Research Digest··2 min read
The authors introduce ScribbleEdit, a benchmark for image editing in which the user's only input is a scribble drawn on the image.

The authors constructed ScribbleEdit with an automated pipeline that generates scribble-image pairs covering distinct editing operations, such as adding, removing or altering objects and attributes.

Why this paper

From ByteDance and 2 others

In one line

Existing VLM/LLM-based image editing models fail to interpret scribble-only editing intents; a soft-token baseline improves understanding.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe