The authors trained Cosmos 3 video world models on nested datasets ranging from 300 to 30,000 hours, a 100-fold scale increase.
More Egocentric Video Improves Agents but Not Their Object Effects
Across training sets from 300 to 30,000 hours, agent motion became realistic much faster than the resulting object dynamics.
Big Tech
Jiahua Dong · Anurag Bagchi · Yash Jangir · Muhammad Zubair Irshad · Sergey Zakharov · Martial Hebert · +4 more
University of Illinois Urbana-Champaign · Carnegie Mellon University · Johns Hopkins University · Toyota Research Institute
Research Digest··3 min read
Thread:Robot World Models
Dong and colleagues trained video world models on progressively larger subsets of a 30,000-hour egocentric manipulation corpus, then separately measured how faithfully the models reproduced human motion and its effects on objects.
Why this paper
From Toyota Research Institute and 3 others · Part of Robot World Models, now 21 papers
In one line
Scaling ego-centric video to 30,000 hours models agents well but leaves object interaction fidelity far behind, requiring training changes, not just more data.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§
What researchers are saying

What 30,000 hours of ego-centric video does not teach - Conditioning saturates the agent with far less data. - With the agent saturated, object interaction converges far below it. - Neither more data nor a larger model closes the gap. arxiv.org/abs/2610.12464
Posts on X from people the desk follows. Their words, unedited.