The author assembled 109 generated two-event clips with manual labels, including near misses where one requested event was visible but the other was absent.
Execution Logs Can Mislead Visual Judges in Video-Generation Agents
Holding video frames fixed, the study shows that auxiliary text can sharply distort some multimodal models’ assessments of whether requested events occurred.
Research Lab
Jian Xu
RIKEN
Research Digest··3 min read
Xu tests whether video-generation verifiers judge visual evidence independently from plans, narration and execution logs.
Why this paper
From RIKEN
In one line
Execution traces cause video-generation verifiers to judge clips based on text, not visuals.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§