Image-based visual examples enable powerful video editing without large paired video datasets

The VINCIE-NExT framework breaks video editing into an image-to-image step followed by image-to-video synthesis, using the edited image as a blueprint for all frames, achieving state-of-the-art results on a benchmark.

Chinese Tech
Leigang Qu · Feng Cheng · Ziyan Yang · Bangbang Yang · Zhaoyang Huang · Wei Chow · +4 more

National University of Singapore · ByteDance Seed · University of Science and Technology of China

Research Digest··3 min read
The authors propose VINCIE-NExT, a video editing framework that leverages in-context visual demonstrations from image editing pairs to guide video edits.

Qu et al.

Why this paper

From ByteDance Seed and 2 others

In one line

Video editing can be achieved by routing edits through the image domain with an in-context visual demonstration, reducing the need for expensive paired video data.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe