The authors constructed FlexBench, a multi-shot benchmark covering 3,105 video shots and 18,161 evaluation queries, with human-verified cross-shot identities and reference-derived checklists for each person's limb actions, stationary limb states, and visibility.
Graded preference optimization improves limb-motion caption fidelity
The authors introduce FlexBench, GPA, and GM-DPO to systematically evaluate and train VLMs on fine-grained limb actions across video shots.
Chinese Tech
Yanan Wang · Tingsong Li · Kaixun Jiang · Chongyang Zhong · Chenwei Xoe · Zhaohe Liao
Zhejiang University · Alibaba Group · University of Science and Technology of China · Fudan University · Shanghai Jiao Tong University
Research Digest··3 min read
The authors present a framework for improving fine-grained limb-motion captioning in Vision-Language Models (VLMs).
Why this paper
From Alibaba Group and 4 others
In one line
Graded preference optimization reduces limb-motion captioning hallucinations by 21.3% over DPO.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§