Benchmark exposes distinct weaknesses in multimodal audio-video generation systems

ORAV Bench tests whether models can selectively combine identities, motion, scenes, speech and other properties from multiple references.

Chinese Tech
Jiacheng Hua · Xiaokun Feng · Jiaqi Hua · Chang Liu · Biao Wang · Miao Liu

Tsinghua University · Tencent Hy

Research Digest··3 min read
The authors introduce ORAV Bench, a benchmark for generating audio and video from several heterogeneous references linked by a textual instruction.

ORAV Bench contains 380 generation tasks, each supplying between 2 and 10 image, video or audio references.

Why this paper

From Tencent Hy and Tsinghua University

In one line

ORAV is a benchmark for audio-video generation from multimodal contexts that reveals distinct model strengths and failures.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.