PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment
Organizations: OriginAI, Israel · NVIDIA
Abstract
Streaming head-avatar reenactment aims to animate a reference image according to a live driving video, requiring robust motion transfer, long-term identity stability, and low latency. Existing methods often rely on specialized identity or motion representations, which can discard useful visual information and inherit failure modes from external extractors. In addition, many recent diffusion-based reenactment methods use offline, clip-based generation, jointly processing and denoising an entire video clip before producing its output, making continuous low-latency streaming difficult. We introduce PixReenact, a pixel-conditioned streaming reenactment framework built on causal video diffusion. PixReenact conditions directly on VAE-encoded reference and driving frames, without specialized identity or motion representations. To separate reference identity from driver motion, we train with cross-identity pseudo supervision together with corrective objectives anchored to the original reference and driving inputs. Long self-rollouts reduce autoregressive drift, while state-aware dual-teacher distillation separately addresses cold-start and steady-state generation. Across three cross-identity benchmarks and a long-horizon streaming benchmark, PixReenact demonstrates robust cross-identity reenactment, particularly under challenging conditions such as extreme viewpoints, occlusions, and pronounced facial expressions, while maintaining the reference identity over long streams. A 4-NFE rolling student continuously emits four frames per update with a mean emission latency of 239 ms.
Figures & tables
| Method | ID | AED | APD | tLP | FVD |
|---|---|---|---|---|---|
| PersonaLive | 0.472 0.23 | 1.295 0.64 | 0.135 0.10 | 1.866 1.57 | 643.5 |
| PixReenact | 0.727 0.11 | 1.135 0.30 | 0.169 0.12 | 1.853 1.35 | 662.5 |
| Public benchmark (75 pairs) | NeRSemble (70 pairs) | Runtime | |||||||||||
| Method | ID | AED | APD | tLP | FVD | ID | AED | APD | tLP | FVD | Lat. | FPS | |
| Per-frame | LivePortrait | 0.863 | 1.071 | 0.208 | 1.146 | 579.0 | 0.849 | 0.787 | 0.0840 | 0.763 | 302.4 | – | 18.90 |
| X-Portrait | 0.701 | 0.887 | 0.097 | 1.613 | 595.2 | 0.774 | 0.694 | 0.0461 | 0.795 | 281.2 | – | 0.65 | |
| Offline | MegActor- | 0.554 | 0.841 | 0.106 | 1.910 | 582.7 | 0.729 | 0.802 | 0.0566 | 0.976 | 247.4 | – | 0.49 |
| Wan-Animate | 0.570 | 0.966 | 0.089 | 1.903 | 579.2 | 0.632 | 0.789 | 0.0381 | 1.199 | 201.3 | – | 0.41 | |
| X-NeMo | 0.614 | 0.775 | 0.065 | 1.378 | 449.1 | 0.710 | 0.639 | 0.0291 | 0.733 | 276.5 | – | 1.37 | |
| Variant | ID | IDD | AED | APD | tLP | FVD | |
| Full recipe (continuation baseline) | 0.707 | 0.023 | 0.816 | 0.071 | 1.254 | 389 | |
| Pixel cond. | Driver attention † | 0.585 | 0.210 | 0.679 | 0.067 | 1.690 | 273 |
| Cross-ID | w/o cross-identity supervision | 0.768 | 0.020 | 0.852 | 0.087 | 1.156 | 458 |
| w/o identity losses † | 0.353 | 0.245 | 0.629 | 0.057 | 1.638 | 293 | |
| w/o motion losses | 0.795 | 0.011 | 0.901 | 0.090 | 1.171 | 548 | |
| State-aware | w/o cold-start teacher | 0.721 | 0.028 | 0.834 | 0.079 | 1.188 | 406 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| wan-labels | PL-labels | |||
|---|---|---|---|---|
| Source | scored | kept | scored | kept |
| NeRSemble | 4,020 | 3,182 | 4,012 | 3,061 |
| VFHQ | 1,450 | 1,106 | 1,447 | 1,084 |
| DH-FaceVid-1K | 10,000 | 7,990 | 9,994 | 7,686 |
| FFHQ | 2,400 | 1,283 | 2,400 | 1,642 |
| Total | 17,870 | 13,561 | 17,853 | 13,473 |
| Term | Definition | Weight | Applies to |
|---|---|---|---|
| surrogate of Eq. 2 (main paper) | 1.0 ( cold) | all | |
| to target latent | 0.3 | all | |
| 1.0 | all a | ||
| 1.0 | cross | ||
| DECA global rotation vs. driver | 1.2 | all | |
| DECA expression+jaw vs. driver | 0.8 | all |
| Pipeline | Stage | Lat. |
| PixReenact | Window assembly (ctx + cond concat) | 0.1 0.0 |
| Denoising DiT forward (4-level staircase, 1 pass) | 155.1 0.2 | |
| Flow + re-noise (scheduler) | 0.7 0.1 | |
| Streaming VAE decode (4 frames, cached) | 82.9 0.5 | |
| Buffer bookkeeping (shift context/buffers) | 0.1 0.0 | |
| Total emission latency | 238.9 0.7 |
| Method / NFE | ID | AED | APD | tLP | FVD | Lat. | FPS |
|---|---|---|---|---|---|---|---|
| PersonaLive | 0.472 | 1.295 | 0.135 | 1.866 | 643.5 | 211.1 | 18.95 |
| PixReenact, 4 (1000,750,500,250) | 0.727 | 1.135 | 0.169 | 1.853 | 662.5 | 238.9 | 16.74 |
| PixReenact, 3 (1000,500,250) | 0.731 | 1.145 | 0.174 | 2.008 | 682.9 | 211.4 | 18.92 |
| PixReenact, 2 (1000,250) | 0.770 | 1.191 | 0.187 | 1.986 | 753.6 | 186.6 | 21.44 |
| Group | CelebV-HQ clip IDs |
|---|---|
| Extreme yaw (25) | -3xE76Q0yLs_6_3, G4f_opQO7MA_4_0, I66kXy4rvME_3, Lrpy2CCnJR0_17, NkjnK6aSrqM_4, QgedeMGTh5E_9, RyevycwXzn4_7, SnBObwk5JF0_37, VCFAT8Hj7Qo_0_0, WGvhuhHYhiA_14, YXetIu8JOzI_3, cWTK6ttZUw0_7, dhSpE_2OF40_9_0, ecsatgCN6Y8_1, hGsR1fkUJqY_10, hOB4Qm1IiOY_0, i1UEj_6T1RE_1, iViIy5ci5yU_24, mpVMM7eqTKs_8_0, sb6b4r8no1w_14_1, t7BPFTnDsJc_1, x99FBmR6OrI_0_0, xGoRE2NsRGI_9, yprYw3FQUpQ_0, zW2T9TkSIcM_5 |
| Extreme pitch (17) | -kFJKQvCrkY_1, 7NHQrVt1yOU_2, 8w0b64kOPss_0, 9DTvF8iBli8_1_3, AcvWJ8bgA8w_3, Fn40RI8ving_1, Jipg0KQ4id0_3, N4XB-3O7UTk_0_0, PFBrK7pbqCU_18, RLbry-3z8yQ_2, Yg6sZ2htZfc_1, dildzih2RPI_3_0, jeycN3BAkt8_38_0, mLh4ZKG50mk_0_0, od6IxUWPMcs_0, qAHkYcmcFfc_6, tnxXQgmHwrk_2 |
| Accessory occlusion (16) | 7FyjCUDR0IM_17, CPKMurScLSE_15_0, Fn40RI8ving_7, HK6TmhLEd2E_12, LjJ276Dh1jA_44, Tn0Ge0GIoGA_10, U8PnHiRRg-s_6, XDmfBTgIRUg_12, XDmfBTgIRUg_15, _wiIxAcoTF4_4, i4ivf-ZmorE_28, qs7FL5DVDMA_10, unmu4yKfBg0_9, xuphB-zbmPk_26, yLElO7gxXPo_43, z7VS7DVHTbc_1_0 |
| Hand occlusion (10) | 2Ke4u7vbLE0_0, M9FAInQwCm0_3, NgtK-eesI7g_0, PP9l4LP0WPI_4, Q9B0pFDCviw_102, dfrJhivMJJY_0, hSiFAK-p-Lg_20, v60fsgGSoDU_0, xu_UgfGriXA_3, yth5JoZKnGQ_1 |
| Benchmark | Pseudo targets | ID | AED | APD | tLP | FVD |
|---|---|---|---|---|---|---|
| Public (75) | Wan-Animate | 0.709 | 0.790 | 0.066 | 1.234 | 377.1 |
| PersonaLive | 0.704 | 0.777 | 0.070 | 1.126 | 402.9 | |
| NeRSemble (70) | Wan-Animate | 0.753 | 0.714 | 0.0426 | 0.803 | 264.8 |
| PersonaLive | 0.735 | 0.689 | 0.0386 | 0.786 | 250.8 | |
| Extreme (68) | Wan-Animate | 0.727 | 1.135 | 0.169 | 1.853 | 662.5 |
| PersonaLive | 0.729 | 1.139 | 0.175 | 1.876 | 653.7 |
| Method | PSNR | SSIM | LPIPS | L1 | ID | AED | APD | tLP | FVD | Lat. | FPS | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TalkingHead-1KH (25 clips) Wang et al. (2021) | ||||||||||||
| VAE round-trip | 38.25 | 0.971 | 0.015 | 0.009 | 0.857 | 0.084 | 0.0019 | 0.184 | 5.7 | – | – | |
| Per-frame | LivePortrait | 19.90 | 0.717 | 0.238 | 0.069 | 0.881 | 0.371 | 0.0256 | 0.856 | 154.2 | – | 18.90 |
| X-Portrait | 20.32 | 0.722 | 0.203 | 0.063 | 0.839 | 0.366 | 0.0226 | 0.832 | 209.1 | – | 0.65 | |
| Offline | MegActor- | 17.47 | 0.664 | 0.284 | 0.086 | 0.825 | 0.460 | 0.0471 | 1.138 | 119.0 | – | 0.49 |
| Wan-Animate | 21.22 | 0.750 | 0.184 | 0.058 | 0.763 | 0.393 | 0.0159 | 1.530 | 106.4 | – | 0.41 | |
| Method | PSNR | SSIM | LPIPS | L1 | ID | AED | APD | tLP | FVD | |
|---|---|---|---|---|---|---|---|---|---|---|
| VAE round-trip | 38.19 | 0.959 | 0.023 | 0.007 | 0.815 | 0.111 | 0.0024 | 0.259 | 8.8 | |
| Per-frame | LivePortrait | 25.88 | 0.810 | 0.168 | 0.030 | 0.848 | 0.341 | 0.0105 | 0.943 | 168.4 |
| X-Portrait | 24.31 | 0.739 | 0.157 | 0.035 | 0.822 | 0.430 | 0.0168 | 0.841 | 139.1 | |
| Offline | MegActor- | 20.07 | 0.704 | 0.279 | 0.062 | 0.743 | 0.642 | 0.0297 | 1.190 | 319.3 |
| Wan-Animate | 24.99 | 0.797 | 0.161 | 0.033 | 0.752 | 0.510 | 0.0151 | 1.679 | 209.3 | |
| X-NeMo | 23.36 | 0.765 | 0.186 | 0.043 | 0.784 | 0.368 | 0.0136 | 0.842 | 133.8 |
| Method | ID | AED | APD | tLP | FVD | |
|---|---|---|---|---|---|---|
| Per-frame | LivePortrait | 0.892 | 1.075 | 0.207 | 1.151 | 636.6 |
| X-Portrait | 0.716 | 0.911 | 0.106 | 1.546 | 639.5 | |
| Offline | MegActor- | 0.570 | 0.829 | 0.100 | 1.911 | 620.0 |
| Wan-Animate | 0.586 | 0.935 | 0.085 | 1.609 | 690.0 | |
| X-NeMo | 0.619 | 0.765 | 0.064 | 1.373 | 510.8 | |
| Wan-Animate-2 | 0.655 | 0.865 | 0.090 | 3.633 | 690.7 |
| Method | ID | AED | APD | tLP |
|---|---|---|---|---|
| Extreme yaw (25 clips) | ||||
| PersonaLive | 0.334 0.20 | 1.416 0.46 | 0.178 0.11 | 2.482 2.20 |
| PixReenact | 0.708 0.09 | 1.302 0.31 | 0.251 0.12 | 2.289 1.89 |
| Extreme pitch (17 clips) | ||||
| PersonaLive | 0.566 0.18 | 1.342 0.64 | 0.115 0.07 | 1.661 1.00 |
| PixReenact | 0.733 0.08 | 1.057 0.15 | 0.126 0.06 | 1.954 0.90 |
| Variant | ID | IDD | AED | APD | tLP | FVD | |
|---|---|---|---|---|---|---|---|
| DMD | with DMD | 0.618 | 0.029 | 0.944 | 0.155 | 0.967 | 508 |
| w/o DMD | 0.643 | 0.026 | 0.966 | 0.157 | 1.178 | 553 | |
| Filtering | filtered | 0.841 | 0.018 | 0.977 | 0.161 | 0.995 | 630 |
| unfiltered | 0.819 | 0.019 | 0.997 | 0.167 | 1.001 | 609 |