The authors designed TriO, a tri-modal unsupervised occupancy world model.
Self-supervised tri-modal model perceives and forecasts any road obstacle without labels
TriO uses camera, LiDAR and RADAR as both inputs and self-supervision to predict 4D occupancy, obstacle segmentation, flow and future LiDAR, achieving state-of-the-art zero-shot obstacle segmentation on Argoverse 2 and Spotting the Unexpected.
Top University
Quinlan Sykora · Sourav Biswas · Christopher Diehl · Andrew Cunningham · Thomas Gilles · Raquel Urtasun
Waabi · University of Toronto
Research Digest··2 min read
The authors present TriO, an unsupervised world model for self-driving perception that uses camera, LiDAR and RADAR as both inputs and self-supervision, removing the need for human annotations.
Why this paper
From University of Toronto and Waabi
In one line
TriO is a tri-modal unsupervised world model that predicts 4D occupancy, obstacle segmentation, flow, and LiDAR, achieving state-of-the-art on multiple 3D/4D tasks and zero-shot road obstacle segmentation.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§