58 million per-second, per-person annotations.
Video-language models often bind actions to the wrong people
ActionLens tests five forms of spatial-temporal reasoning and finds a large gap between leading models and human performance.
Big Tech
Gueter Josmy Faure · Min-Hung Chen · Hao Ping Wang · Timothée Lardy · Hung-Ting Su · Winston H. Hsu
National Taiwan University · NVIDIA
Research Digest··2 min read
Faure and colleagues introduce ActionLens, a 6,701-question video benchmark designed to test whether vision-language models can associate an action with the correct person and moment.
Why this paper
From NVIDIA and National Taiwan University
In one line
Video-capable vision-language models fail to associate the right action with the right person at the right moment.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§