VidScribe covers four regimes: generating written text from a language prompt, transferring text from a visual reference, preserving text during image-to-video generation, and editing text within an existing video.
Video generators preserve existing text better than they edit it
VidScribe tests text rendering across four video-generation tasks and finds that readable content, correct glyph shapes and reliable editing are distinct capabilities.
Chinese Tech
Ziying Zhang · Litao Li · Junchao Liao · Tianyi Zeng · Siyu Zhu · Long Qin · +1 more
Alibaba Group · Shanghai Jiao Tong University · Fudan University
Research Digest··3 min read
Zhang and colleagues introduce VidScribe, a benchmark of 803 human-verified samples designed to expose failures in generating, preserving and editing text inside video.
Why this paper
From Alibaba Group and 2 others · Released data
In one line
Video text generation capability is non-monolithic: I2V sustains text best, V2V editing is weakest, and degradation concentrates on structural and temporal factors.
What it released
Data
What we could check
- ·No code link found
- ·No weights link found
- ✓Dataset link in the paper (huggingface.co)
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§