New benchmark reveals streaming VLMs struggle with high-dynamic video

FastBench tests models on fast-paced real-world streams, showing dense sampling helps but still falls short of oracle-guided perception.

Chinese Tech
Yuxuan Hu · Weikang Shi · Yang Bo · Xudong Lu · Xintong Guo · Shuhan Li · +7 more

CUHK MMLab · Huawei Research

Research Digest··2 min read
The authors introduce FastBench, a benchmark of 306 QA pairs across eight domains designed to evaluate streaming video language models on high-dynamic content.

The authors built FastBench using a trajectory-grounded pipeline: QA generation from high-FPS clips, filtering out questions answerable at 2 FPS, verification with SAM3 and CoTracker3 trajectories, and three rounds of human inspection.

Why this paper

From Huawei Research and CUHK MMLab

In one line

Streaming VLMs score at most 50.7% on high-dynamic real-world streams in FastBench, and denser sampling yields only partial gains.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe