Abstract
In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance while enabling flexible visual edits remains challenging for generative video models. We formalize this challenge as identity-preserving video restylization, which propagates scene, lighting, and style changes specified by an edited keyframe across a source video, while preserving facial likeness and performance, including expressions, eye gaze, and lip synchronization. A key obstacle is the absence of paired training data, as identity-preserving restylized video pairs are rare in real-world settings. To address this, we propose a decoupling of source-grounded identity preservation and edit-driven video synthesis. Our key insight is that facial appearance and expression should remain invariant, with illumination being the primary permissible variation. We therefore cast identity preservation as a video relighting problem, while modeling visual edit propagation as controlled video synthesis guided by the edited keyframe. Building on this formulation, we introduce ID-V2V, a video-to-video generative framework integrating complementary control signals: relit facial regions and facial normal maps tightly constrain facial likeness and performance, while edited keyframes and depth sequences enable flexible and temporally coherent generation. This design enables constructing training pairs from a single video, eliminating the need for scarce paired data. Extensive experiments demonstrate that ID-V2V significantly outperforms existing methods in preserving facial likeness and fine-grained facial performance, supports both single- and multi-subject scenarios, and delivers high visual quality, highlighting its potential as a human-centric tool for real-world content production. The code is available at: https://github.com/Eyeline-Labs/ID-V2V.
Community
Capture the performance first. Redesign the look later.
Introducing ID-V2V, Netflix’s latest research exploration in human-centric video editing, to appear at SIGGRAPH Asia 2026. It enables a powerful creative workflow: creators can redesign the environment, lighting, and visual style of an existing video after capture while preserving the original human identity and performance.
Built for production workflows, ID-V2V introduces identity-preserving video restylization. Given a source video and edited keyframes, ID-V2V propagates those edits across the entire video while preserving human appearance, subtle facial expressions, full-body motion, and multi-person interactions. This gives creators the freedom to reshape the visual world in post-production while keeping the original performance intact.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Vera: Identity-Faithful Human Subject-to-Video Generation (2026)
- Customizing Video Portraits via Identity-ActionDecoupling (2026)
- ViDS: Video Diffusion Shader using 3D Face Tracking (2026)
- TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration (2026)
- TriMotion: Modality-Agnostic Camera Control for Video Generation (2026)
- Geometry-Instructed Video Editing (2026)
- Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.22830 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper