Language-Based Memory Helps Driving Models Track Earlier Road Events

AD-Memo records driving-relevant objects in text, reuses those records in later decisions, and can share them with other models.

Big Tech
Kai Yan · Xiangyu Chen · Yulong Cao · Alex Naumann · Peter Karkus · Yan Wang · +6 more

University of Illinois Urbana-Champaign · NVIDIA

Research Digest··3 min read
Yan and colleagues introduce AD-Memo, a vision-language-action driving agent that converts selected observations into textual memory rather than repeatedly processing a long sequence of camera frames.

AD-Memo extends the model's language reasoning output with a memory describing surrounding objects important for future driving.

Why this paper

From NVIDIA and University of Illinois Urbana-Champaign

In one line

AD-Memo gives autonomous-driving VLA agents reusable language memory, improving memory-dependent driving, scene question answering, and cross-model portability.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.