The authors built CinematicVQA from the Video Aesthetic Database, which contains expert commentary on video segments.
Video models struggle to connect cinematic techniques with storytelling intent
CinematicVQA tests whether vision-language models can trace composition, lighting and camera choices to perceptual effects and narrative functions.
Big Tech
Shuo Xing · Pooja Verlani · Balu Adsumilli · Zhengzhong Tu
Texas A&M University · Google Inc.
Research Digest··2 min read
Xing and colleagues introduce CinematicVQA, a benchmark for reasoning about how cinematographic choices shape what viewers perceive and how a scene functions narratively.
Why this paper
From Google Inc. and Texas A&M University
In one line
Large vision-language models can describe film visuals but fail to identify the techniques that produce them.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§