More Egocentric Video Improves Agents but Not Their Object Effects

Across training sets from 300 to 30,000 hours, agent motion became realistic much faster than the resulting object dynamics.

Big Tech
Jiahua Dong · Anurag Bagchi · Yash Jangir · Muhammad Zubair Irshad · Sergey Zakharov · Martial Hebert · +4 more

University of Illinois Urbana-Champaign · Carnegie Mellon University · Johns Hopkins University · Toyota Research Institute

Research Digest··3 min read
Dong and colleagues trained video world models on progressively larger subsets of a 30,000-hour egocentric manipulation corpus, then separately measured how faithfully the models reproduced human motion and its effects on objects.

The authors trained Cosmos 3 video world models on nested datasets ranging from 300 to 30,000 hours, a 100-fold scale increase.

Why this paper

From Toyota Research Institute and 3 others · Part of Robot World Models, now 21 papers

In one line

Scaling ego-centric video to 30,000 hours models agents well but leaves object interaction fidelity far behind, requiring training changes, not just more data.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

What researchers are saying

Aran Komatsuzaki @arankomatsuzaki
AI research curator
What 30,000 hours of ego-centric video does not teach - Conditioning saturates the agent with far less data. - With the agent saturated, object interaction converges far below it. - Neither more data nor a larger model closes the gap. arxiv.org/abs/2610.12464

Posts on X from people the desk follows. Their words, unedited.

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe