The authors formalize the task of collection-grounded B-roll generation: given a corpus of user footage, a textual directive, and a desired length, produce a coherent multi-shot video sequence that complements the primary footage while preserving the visual world of the collection.
Generating B-roll footage grounded in a user's captured video collection.
A three-stage system called MemComposer constructs an entity-indexed memory from raw footage, then plans, retrieves, and iteratively generates coherent B-roll sequences.
Big Tech
Cusuh Ham · Fabian Caba Heilbron · Josef Sivic · Bryan Russell
Adobe Research · Czech Technical University
Research Digest··3 min read
The authors introduce MemComposer for collection-grounded B-roll generation.
Why this paper
From Adobe Research and Czech Technical University
In one line
MemComposer generates B-roll sequences from user video collections that are grounded in the visual world and adhere to natural language directives.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§