The authors adapted a compact video-diffusion backbone for causal, autoregressive generation, meaning each output is produced from earlier states rather than by denoising a complete video clip.
Pixel conditioning enables stable, low-latency streaming head-avatar reenactment
PixReenact transfers motion directly from driving pixels while using causal diffusion to generate an avatar continuously.
Big Tech
Gavriel Habib · Dvir Samuel · Or Shimshi · Rami Ben-Ari
OriginAI · NVIDIA
Research Digest··2 min read
Habib et al.
Why this paper
From NVIDIA and OriginAI
In one line
PixReenact uses causal, pixel-conditioned video diffusion to stream identity-stable head-avatar reenactment, emitting four frames per update at 239 ms mean latency.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§