Video generators preserve existing text better than they edit it

VidScribe tests text rendering across four video-generation tasks and finds that readable content, correct glyph shapes and reliable editing are distinct capabilities.

Chinese Tech
Ziying Zhang · Litao Li · Junchao Liao · Tianyi Zeng · Siyu Zhu · Long Qin · +1 more

Alibaba Group · Shanghai Jiao Tong University · Fudan University

Research Digest··3 min read
Zhang and colleagues introduce VidScribe, a benchmark of 803 human-verified samples designed to expose failures in generating, preserving and editing text inside video.

VidScribe covers four regimes: generating written text from a language prompt, transferring text from a visual reference, preserving text during image-to-video generation, and editing text within an existing video.

Why this paper

From Alibaba Group and 2 others · Released data

In one line

Video text generation capability is non-monolithic: I2V sustains text best, V2V editing is weakest, and degradation concentrates on structural and temporal factors.

What it released

Data

What we could check

  • ·No code link found
  • ·No weights link found
  • ✓Dataset link in the paper (huggingface.co)
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.