The authors introduce Behavior Pack Optimization (BPO), a reinforcement-learning framework for video multimodal large language models.
Training video models across counterfactual views improves evidence-grounded answers
Behavior Pack Optimization jointly rewards correct, stable, intervention-sensitive and appropriately uncertain responses across altered versions of a video.
Top University
Zhaolu Kang · Shiyu Liu · Tailong Luo · Wei Zhang · Yingjie He · Lei Wei · +9 more
Peking University · The University of Melbourne · Stanford University · Chengdu Minto Tech
Research Digest··2 min read
Kang and colleagues post-trained video question-answering models on packs containing an original clip and targeted counterfactual versions, rather than grading each response independently.
Why this paper
From Peking University and 3 others
In one line
Optimizing video MLLMs on behavior packs across counterfactual views improves accuracy, temporal reasoning, and abstention over single-response post-training.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§