Grounding long-video memories in visual identity lets QA systems track objects across days

The authors' GEB framework links observations of the same physical object across clips and retrieves those linked biographies when answering questions.

Big Tech
Hui Ren · Lei Fan · Henry Pao · Han Guo · Zeeshan Zia · Ying Chen · +2 more

University of Illinois Urbana-Champaign · Amazon.com, Inc.

Research Digest··2 min read
The authors introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies.

GEB addresses a failure mode of chronological or text-derived video memory: descriptions can conflate distinct objects that share a name, and fragment a single object observed under different states or locations.

Why this paper

From Amazon.com, Inc. and University of Illinois Urbana-Champaign

In one line

Grounded Entity Biographies link visually identified objects across video clips to answer questions about long recordings.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.