Video virtual try-on has attracted increasing attention due to its broad potential in digital fashion and intelligent e-commerce. However, existing methods primarily focus on low-resolution settings and still face substantial challenges when extended to high-resolution scenarios. These limitations can be attributed to two main factors: (1) the insufficient utilization of rich garment reference information, and (2) the lack of explicit positional modeling between garment and video representations during cross-modal interaction, which weakens fine-grained local correspondence. To address these issues, we propose TexTailor, a high-fidelity video virtual try-on framework built upon a pretrained video Diffusion Transformer. Specifically, we introduce a timestep-adaptive modulation mechanism to dynamically adjust garment visual representations throughout denoising. We further develop a frame-aligned positional encoding strategy to strengthen garment-to-video correspondence, together with a multi-source injection design that reduces interference among heterogeneous conditions. Extensive experiments on multiple video virtual try-on benchmarks, including the high-resolution Eevee dataset, demonstrate that TexTailor achieves competitive performance in garment detail preservation, temporal consistency, and overall video quality.
Figures & tables
Figure 1: TexTailor generates high-fidelity video virtual try-on results with temporally consistent motion while preserving garment textures, patterns, and local structures across diverse poses and viewpoints.
Figure 2: Overview of TexTailor. (a) A DiT backbone integrates structural input, text token, visual token, and garment latent token through MCAI for video try-on generation. (b) TAVM modulates visual features with timestep-adaptive token-wise gating. (c) FAC-RoPE assigns frame-aligned temporal coordinates to static garment latent tokens, preserving native 3D RoPE while establishing stable garment-to-video spatial correspondence.
Figure 3: Qualitative comparison on the high-resolution Eevee benchmark. Compared with existing methods, TexTailor better preserves garment textures, patterns, and local structures while maintaining temporally coherent appearance across frames.
Method
Full-shot
Close-up
GPU Mem.
Time
VFID _R↓
VFID _I↓
VGID ↑
VFID _R↓
VFID _I↓
VGID ↑
ViViD ( Fang et al. 2024 )
0.565
12.859
0.514
1.253
12.665
0.533
72.21G
348.96s
MagicTryOn ( Li et al. 2025b )
0.187
9.783
0.512
0.752
11.285
0.538
69.54G
589.73s
CatV 2 TON ( Chong et al. 2025 )
0.751
9.086
0.520
0.673
11.935
0.522
40.18G
357.21s
TexTailor
0.149
8.872
0.531
0.489
10.917
0.543
30.32G
230.72s
Table 1: Quantitative comparison on the Eevee benchmark at ( 1088×816 ). We report results under both full-shot and close-up settings, together with GPU memory usage and inference time. The best and second-best results are highlighted in bold and underline , respectively.
Method
VFID Ip↓
VFID Rp↓
SSIM ↑
LPIPS ↓
VFID Iu↓
VFID Ru↓
GPU Mem.
Time
ViViD ( Fang et al. 2024 )
17.1847
0.6382
0.8041
0.1216
21.6925
0.8347
62.59G
204.183s
CatV 2 TON ( Chong et al. 2025 )
13.4821
0.2876
0.8742
0.0647
19.3846
0.5179
27.66G
209.127s
MagicTryOn ( Li et al. 2025b )
8.3165
0.2298
0.9024
0.0598
14.6032
0.3137
51.51G
345.271s
TexTailor
7.9428
0.2645
0.9017
0.0619
13.8754
0.2861
25.32G
196.852s
Table 2: Quantitative comparison on the ViViD benchmark. We report paired and unpaired evaluation metrics, together with GPU memory usage and inference time. The best and second-best results are highlighted in bold and underline , respectively.
Setting
Metric
w/o TAVM-S
w/o TAVM-D
w/o TAVM-G
w/o TAVM
w/o FAC
w/o MCAI-T
w/o MCAI-G
w/o Mask
Full Model
Full-shot
VFID R ↓
0.184
0.181
0.176
0.224
0.207
0.179
0.201
0.198
0.149
VFID I ↓
9.512
9.463
9.382
10.024
9.781
9.427
9.694
9.821
8.872
VGID ↑
0.520
0.521
0.522
0.512
0.515
0.523
0.516
0.517
0.531
Close-up
VFID R ↓
0.551
0.543
0.535
0.612
0.581
0.529
0.568
0.592
0.489
VFID I ↓
11.612
11.534
11.462
12.146
11.873
11.421
11.754
11.982
10.917
VGID ↑
0.542
0.535
0.526
0.514
0.539
0.527
0.540
0.538
0.543
Table 3: Ablation study on the Eevee benchmark. We evaluate the contribution of each component in TexTailor.
Figure 4: Visualization of timestep-wise token gating in TAVM. (a) Gate activation maps across different denoising steps. (b) Representative token regions selected from the garment reference. (c) Gate evolution curves of the selected token regions in (b). TAVM progressively shifts its emphasis from coarse garment structure to fine-grained local details during denoising.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Full-shot
Close-up
GPU Mem.
Time
VFID R ↓
VFID I ↓
VGID ↑
VFID R ↓
VFID I ↓
VGID ↑
ViViD ( Fang et al. 2024 )
0.389
12.194
0.506
0.936
12.198
0.533
64.12G
216.734s
MagicTryOn ( Li et al. 2025b )
0.161
9.865
0.520
0.595
11.262
0.534
49.83G
333.906s
CatV 2 TON ( Chong et al. 2025 )
0.746
9.141
0.518
0.632
11.847
0.538
29.14G
197.683s
TexTailor
0.153
8.996
0.529
0.497
11.071
0.553
24.87G
188.436s
Appendix
Table 4: Quantitative comparison on the Eevee benchmark at 832×624 . We report results under both full-shot and close-up settings, together with GPU memory usage and inference time. The best and second-best results are highlighted in bold and underline , respectively.
Figure 5: Qualitative comparison on challenging garment cases. TexTailor better preserves complex patterns, fine-grained textures, and local garment structures compared with existing methods.
Figure 6: Qualitative comparison across diverse garment categories. TexTailor achieves more faithful garment transfer with improved texture preservation and structural consistency.
Figure 7: Additional results of TexTailor on diverse garment categories, including sweaters, dresses, and pants.
Figure 8: Additional results of TexTailor on garments with complex patterns and structures.
Figure 9: Additional results demonstrating the generalization ability of TexTailor across various garments and motion scenarios.
Figure 10: Qualitative ablation comparison of TexTailor. Removing individual components leads to degraded garment fidelity, including texture distortion, blurred details, and weaker pattern preservation. The full model achieves the most faithful garment appearance.
Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches cast VVT as mask-conditioned video inpainting and rely on separate modules for human parsing, pose estimation, and garment warping. This multi-stage design complicates deployment and, more critically, allows errors in explicit geometric priors to propagate irreversibly into the generated video. We present UniVVT, a unified end-to-end framework that reframes VVT as semantically conditioned video generation, eliminating mask, pose, and warping modules at inference. At its core, a scene-task perceiver built on a Multimodal Large Language Model jointly encodes the source video, target garment, and task instruction into compact, task-aware latent tokens, implicitly capturing what to transfer and where and how to transfer it. A lightweight semantic bridge then aligns these tokens with the conditioning space of a diffusion-based video generator, enabling coherent garment transfer. To robustly couple the heterogeneous components, we devise a three-stage progressive training strategy comprising semantic alignment, joint task adaptation, and flexible-resolution refinement. Extensive experiments demonstrate that UniVVT achieves state-of-the-art performance across multiple benchmarks, validating implicit semantic guidance as a simple and effective alternative to fragile geometric preprocessing for end-to-end virtual try-on.
Yushe Cao, Shikun Feng, Fei Shen +5
Tsinghua University · Zhongguancun Academy · National University of Singapore +2
Video Virtual Try-On (VVT) aims to seamlessly replace a garment on a person in a video with a new one. While existing methods have made significant strides in maintaining temporal consistency, they are predominantly confined to non-interactive scenarios where models merely showcase garments. This limitation overlooks a crucial aspect of real-world apparel presentation: active human-garment interaction. To bridge this gap, we introduce and formalize a new challenging task: Interactive Video Virtual Try-On (Interactive VVT), where subjects in the video actively engage with their clothing. This task introduces unique challenges beyond simple texture preservation, including: (1) resolving the semantic ambiguity of interactions from standard pose information, and (2) learning complex garment deformations from video where interactive moments are sparse and brief. To address these challenges, we propose iTryOn, a novel framework built upon a large-scale video diffusion Transformer. iTryOn pioneers a multi-level interaction injection mechanism to guide the generation of complex dynamics. At the spatial level, we introduce a garment-agnostic 3D hand prior to provide fine-grained guidance for precise hand-garment contact, effectively resolving spatial ambiguity. At the semantic level, iTryOn leverages global captions for overall context and time-stamped action captions for localized interactions, synchronized via our novel Action-aware Rotational Position Embedding (A-RoPE). Extensive experiments demonstrate that iTryOn not only achieves state-of-the-art performance on traditional VVT benchmarks but also establishes a commanding lead in the new interactive setting, marking a significant step towards more dynamic and controllable virtual try-on experiences.
Jun Zheng, Zhengze Xu, Mengting Chen +6
Shenzhen Campus of Sun Yat-sen University · Taobao & Tmall Group of Alibaba
Although video virtual try-on (VVT) has achieved significant progress, existing methods still exhibit two fundamental limitations: first, they are restricted to single-garment transfer, rendering simultaneous multi-object try-on highly impractical; second, their heavy reliance on explicit external priors (e.g., garment masks) inevitably destroys crucial physical dynamics and degrades visual quality. To bridge this gap, this paper proposes the novel Try-On Anything task, which aims to simultaneously transfer diverse wearable objects onto a person in a video in a single inference pass. To support and standardize this paradigm, we introduce TryAny-Bench, a comprehensive benchmark encompassing a paired video dataset alongside a tailored evaluation protocol. Furthermore, we present OmniTryOn, an external-prior-free generative framework designed to tackle this task. Specifically, OmniTryOn employs a First Frame Wearable Cache strategy, which directly provides diverse wearable objects for the generation process through the initial video frame. To maintain consistency, we propose the Spatiotemporally Consistent RoPE (STC-RoPE), which inherently establishes robust spatiotemporal anchors to strictly preserve complex human motions and background dynamics. Optimized by the proposed Gradual Try-On (GTO) training strategy, our model progressively masters robust multi-object synthesis. Extensive experiments on TryAny-Bench demonstrate that OmniTryOn significantly outperforms existing specialized video virtual try-on models and general video editing baselines, establishing a powerful new standard for the Try-On Anything task. Our dataset, code, and models are available at https://github.com/xcltql666/OminTryOn.