Self-supervised tri-modal model perceives and forecasts any road obstacle without labels

TriO uses camera, LiDAR and RADAR as both inputs and self-supervision to predict 4D occupancy, obstacle segmentation, flow and future LiDAR, achieving state-of-the-art zero-shot obstacle segmentation on Argoverse 2 and Spotting the Unexpected.

Top University
Quinlan Sykora · Sourav Biswas · Christopher Diehl · Andrew Cunningham · Thomas Gilles · Raquel Urtasun

Waabi · University of Toronto

Research Digest··2 min read
The authors present TriO, an unsupervised world model for self-driving perception that uses camera, LiDAR and RADAR as both inputs and self-supervision, removing the need for human annotations.

The authors designed TriO, a tri-modal unsupervised occupancy world model.

Why this paper

From University of Toronto and Waabi

In one line

TriO is a tri-modal unsupervised world model that predicts 4D occupancy, obstacle segmentation, flow, and LiDAR, achieving state-of-the-art on multiple 3D/4D tasks and zero-shot road obstacle segmentation.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.