4D foundation models lack robust visual memory for out-of-view objects

New PersistBench benchmark uses 360° videos to show that current video and reconstruction models degrade sharply once objects leave the field of view.

Industry
Guangzhao He · Hadar Averbuch-Elor · Wei-Chiu Ma

Cornell University

Research Digest··2 min read
The authors introduce PersistBench, a dataset and metric suite using 360° videos as ground truth to test object permanence, motion continuity, and appearance preservation in 4D foundation models.

The authors developed PersistBench, a benchmark comprising 360° videos that provide omniscient ground truth beyond the traditional field-of-view.

Why this paper

From Cornell University · Part of World Model Planning for Agents, now 19 papers

In one line

Current 4D foundation models fail to remember objects that exit their field of view.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.