The authors built FastBench using a trajectory-grounded pipeline: QA generation from high-FPS clips, filtering out questions answerable at 2 FPS, verification with SAM3 and CoTracker3 trajectories, and three rounds of human inspection.
New benchmark reveals streaming VLMs struggle with high-dynamic video
FastBench tests models on fast-paced real-world streams, showing dense sampling helps but still falls short of oracle-guided perception.
Chinese Tech
Yuxuan Hu · Weikang Shi · Yang Bo · Xudong Lu · Xintong Guo · Shuhan Li · +7 more
CUHK MMLab · Huawei Research
Research Digest··2 min read
The authors introduce FastBench, a benchmark of 306 QA pairs across eight domains designed to evaluate streaming video language models on high-dynamic content.
Why this paper
From Huawei Research and CUHK MMLab
In one line
Streaming VLMs score at most 50.7% on high-dynamic real-world streams in FastBench, and denser sampling yields only partial gains.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§