Vera: A Layered Diffusion Model for Content-Preserving Video Editing
- Type
- paper
- Venue
- arXiv (cs.CV)
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Introduces Vera (from Latin vera, "genuine"), a layered diffusion framework for content-preserving video editing. Instead of regenerating the entire video, Vera generates an edit layer along with an alpha matte for compositing with the source video, separating creative editing from content preservation by design. To encourage coherent composition, it extends the text-to-video DiT into a Mixture-of-Transformers (MoT) with separate DiTs per layer interacting through joint self-attention, and constructs a high-quality layered dataset with accurate alpha mattes, diverse scenes and dynamics, and visual effects.
Research notes
- Key findings: Content preservation by design: generates an edit layer plus an alpha matte composited with the source video instead of regenerating every pixel; Mixture-of-Transformers (MoT) architecture: separate DiTs per layer that interact through joint self-attention for coherent composition with the source video; Outperforms leading open-source video editing models in content preservation while remaining competitive in edit quality, per a quantitative benchmark and human-preference study, using 486K frames of layered training data
- Work done during an internship at Netflix; project lead Zhuoning Yuan (zyuan@netflix.com). Project page: https://veralayereddiffusion.github.io/. Associated dataset: netflix/Vera-Layered-Video-Dataset (same batch).