The authors built PhysWAM around a single flow-matching transformer, a generative model trained to turn noise into structured outputs.
Shared 3D geometry aligns predicted driving scenes and vehicle motion
PhysWAM jointly generates future video, metric depth and ego trajectories, tying them together through a LiDAR-based geometric training constraint.
Big Tech
Dhruv Parikh · Fengcheng Yu · Quankai Gao · Jiawei Yang · Junjie Ye · Maulik Bhatt · +8 more
University of Southern California · Woven by Toyota · Toyota Research Institute · DEVCOM Army Research Office
Research Digest··2 min read
Parikh et al.
Why this paper
From Toyota Research Institute and 3 others
In one line
PhysWAM uses a geometric constraint between depth and ego motion to jointly predict future scenes and actions for autonomous driving.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§