Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. These capabilities include single-pass classifier-free guidance, few-step sampling, and long-context error correction for autoregressive generation. The adapters remain reusable even when downstream models add conditioning branches, expand output channels. Despite training at a fixed guidance scale, our dedicated CFG LoRA provides text guidance control through its inference weight. Combining it with a few-step LoRA simultaneously preserves few-step generation and CFG controllability on downstream tasks. We verify training-free deployment on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation. The approach may support additional compatible models. Each capability can thus be distilled once per backbone family and reused without per-target retraining.
Figures & tables
Figure 1 : Decoupled guidance control through an additional CFG-only LoRA. (a) Scaling a single distilled LoRA changes the entire update, including its learned guidance and few-step behavior. (b) Adding a separately weighted CFG-only LoRA lets us adjust guidance for each target while keeping the few-step adapter weight fixed. Arrow colors and widths schematically illustrate guidance compatibility, with the warmest coupled-transfer arrow at CFG 5 . CFG-only LoRA scaling provides an adjustable guidance control.
Figure 3 : Matched transfer comparisons on SCOPE and Wan2.2-Fun-5B-Control (Depth). Two matched frames per task compare native (30/40 steps), naive four-step, per-target distilled, and LongLive-Plug outputs. Naive four-step sampling produces blurry, low-quality videos, whereas LongLive-Plug maintains high visual quality at four steps. Depth thumbnails condition ControlNet; boxes and strips show matched regions across methods. LongLive-Plug uses an additional base-distilled adapter variant. Appendix B gives full four-frame comparisons and alignment details.
Figure 4 : Guidance control after CFG-only distillation. The milk-splatter prompt compares native CFG references with CFG LoRA weights 1 , 2 , and 3 . All variants use 50 sampling steps; every LoRA variant uses the distilled CFG setting and requires only one conditional forward pass per step. Rows show matched frames at 1 and 4 seconds. Higher LoRA weights follow the trend of stronger CFG, producing a more pronounced splash without exactly matching native CFG scales.
Figure 5 : Independent CFG control on SCOPE. (A) The coupled LoRA alone at weight 1 shows little response to the prompts. Adding a separately weighted CFG LoRA strengthens the boxed prompt attributes while the coupled LoRA weight remains at 1 . Colored lines link each box to its corresponding prompt text. (B) Scaling the entire coupled LoRA instead degrades generation. Frames are matched across weights within each row. All runs use four-step sampling with distilled CFG; (A) and (B) use different coupled checkpoints.
Figure 6 : Rank, data diversity, and long-context transfer. (a–b) Higher rank and more diverse prompts (lower concentration) improve FVD after transfer. (c–d) Mean of seven VBench dimensions after transfer to ReWorld and Matrix-Game 3.0, respectively. Long-context comparisons are within each model. Our transferred long-context LoRA improves quality during long AR rollouts.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
CFG-only, Wan2.2
Few-step, Wan2.2
Few-step, Wan2.1
LoRA training rank r
128
128
128
Fake-score LoRA
None
Yes
Yes
Generator / critic LR
10−5 / —
10−5 / 2×10−6
10−5 / 2×10−6
AdamW (β1,β2)
(0,0.999)
(0,0.999)
(0,0.999)
Weight decay
0
0
0.01
LoRA dropout
0
0
0
Appendix
Table 4 : Wan training rank and reference optimization settings. All branches are trained at rank 128. Batch sizes are per process. For DMD, the generator updates once per five fake-score updates. The CFG column reports the reference CFG-only optimization settings.
Figure 7 : Keyframe comparison on SCOPE. Four matched keyframes from an 81-frame sequence at 20 fps. Rows compare default 30-step inference, naive four-step sampling, SCOPE-specific distillation, and transferred LoRAs, using the same case and seed. Boxes mark identical image coordinates across methods; the strips below each frame magnify these regions. Red highlights the naive four-step row’s ghosted wall and door edges, while the distilled variants retain clearer boundaries and surface detail.
Figure 8 : Keyframe comparison on Wan2.2-Fun-5B-Control. Four matched keyframes compare input depth and four generation methods. Boxes mark identical image coordinates across methods, magnified in the strips below each output. Red highlights the naive four-step row’s smeared rock, foliage, and road details and low contrast. Columns match conditioning frame indices (depth at 30 fps, outputs at 24 fps). The transferred output uses an additional base-distilled adapter variant.
H3-World [ 10 ] ; Code World Model [ 12 ] ; SolarWM-H3 [ 32 ]
—
Fun ControlNet-Union [ 6 ]
SolarWM-H3 [ 32 ]
Appendix
Table 5 : Recorded transfer coverage across backbone families and tasks. Every listed Wan-model entry is evaluated for four-step and CFG transfer. The two panels share the same backbone axis.
Figure 9 : Additional Wan2.1-14B transfer cases. LongVie 2, ABot-PhysWorld, MagicTryOn, and Wan-Alpha illustrate world modeling, robotics, subject conditioning, and RGBA output. Each row compares two native frames (left) with the same two timestamps after four-step transfer (right). Inputs and task conditions are paired in the source report. The checkerboard is part of the Wan-Alpha preview, not a measurement of alpha-channel accuracy. Changes in appearance and motion remain visible after transfer.
Figure 10 : Additional Wan2.2-TI2V-5B transfer cases. Depth-conditioned Fun Control, FlashMotion, Kiwi-Edit, and Loomis Painter cover structure, trajectory, editing, and style adaptation. The native and transferred columns show identical frame indices for each paired task input. Native inference uses 50 steps for Kiwi-Edit and Loomis Painter. All transferred outputs use four steps with CFG distilled into the adapter; downstream conditioning and task adapters are retained.
Figure 11 : H3 task transfer: action control and line-art coloring. Left: undistilled multi-step inference. Middle: naive four-step sampling. Right: four-step inference with LongLive-Plug. H3-World shows matched frames at 1 and 4 s under a forward-action condition. LineartAnime shows frames at 1 s and the final available frame (3.71 s). S4 preserves a clearer character outline than E4 in these examples, while appearance can differ from D. These static frames do not evaluate audio quality or synchronization.
Figure 12 : H3 camera transfer includes a quality–control trade-off. Left: undistilled multi-step inference. Middle: naive four-step sampling. Right: four-step inference with LongLive-Plug. In the museum-crane example, S4 retains clearer architectural detail and a visible upward-camera response. In the library-yaw example, S4 remains sharp but its change in framing is attenuated relative to D/E4. HUD elements are inherited from the supplied anchor images. These cases have no generated-video pose regression metric and do not establish precise trajectory adherence.
Figure 13 : Additional Wan CFG-only control example. Native CFG references and CFG LoRA weights 1 , 2 , and 3 use 50 sampling steps; LoRA outputs use runtime CFG 1 with one conditional evaluation per step. Matched frames at 1 and 4 s show more pronounced flower opening as the adapter weight increases.
Figure 14 : Native CFG and CFG-only LoRA on Rooftop Martial Arts. Matched frames at 1 and 4 s compare native CFG (top: no CFG and scales 2 – 5 ) with CFG LoRA (bottom: weights 0.5 , 1 , 2 , 3 , 4 ). The prompt, seed, and native schedule are fixed; LoRA outputs use one conditional forward pass per step. LoRA weights are not calibrated native CFG scales.
Figure 15 : Native CFG and CFG-only LoRA on the two night-village cases. Each case places native CFG above CFG LoRA at matched 1 and 4 s frames, using the sweeps in Fig. 14 . The prompt, seed, and native sampling schedule are fixed within each case.
Figure 16 : CFG-branch control with four-step SCOPE generation. The prompts request snow cover in a mountain valley and autumn vegetation around an ancient temple. Both use 81 frames at 20 fps; rows show 1, 2, and 4 s. Increasing the CFG-branch weight strengthens the snow cover and orange-red foliage, while the few-step weight stays at 1 . An excessively large weight ( 5 ) degrades quality, changing geometry and introducing spurious text.
Figure 17 : CFG-branch control with four-step video continuation. We continue the same watercolor paper-boat input with golden butterflies or pink lotus blossoms. The last 24 input frames condition 81 new frames at 24 fps; displayed times are relative to the generated continuation. Weights 2 – 3 produce more visible and persistent butterflies or more prominent lotus blossoms. At weight 5 , dense generated content comes with fragmented scenery and stronger changes to the boat and islands. The few-step branch remains fixed at weight 1 in every column.
Figure 18 : Long-rollout qualitative comparisons on ReWorld. Matched frames from a modern interior and a country lane. The +Long outputs retain more visible texture and object detail at late times.
Figure 19 : Long-context transfer to Matrix-Game 3.0: animated city. Columns show native frames 132, 528, 924, and 1,056 (zero-based), spanning early, middle, and late stages of the same 62.18 s rollout. The +Long output retains distinct facade edges and street objects at late times.
Figure 20 : Long-context transfer to Matrix-Game 3.0: overgrown temple. The same methods and timestamps as Fig. 19 compare stone architecture and vegetation. The +Long frames preserve visible stone-block boundaries, steps, and foliage late in the rollout, while appearance and layout vary across methods.
Recently, diffusion models have made great progress in video generation. However, most existing video diffusion models are trained with short videos, and degrade when extrapolated to long videos, struggling to maintain long-range temporal coherence while retaining diverse motions. To generate consistent, high-quality and dynamic long videos, we propose Diff-VF, a training-free, plug-and-play and model-agnostic framework that converts existing short-video diffusion backbones into long-video generators without modifying or fine-tuning the base model. Diff-VF couples three complementary strategies: Hybrid Noise Initialization (HNI) to constrain global semantics, Weighted Window Sampling (WWS) to remove inter-window discontinuities, and Temporal Extended Sampling (TES) to establish long-range dependencies with a timestep-varying fusion. We further extend Diff-VF to long-video enhancement via Skip Residual Guidance that balances fidelity and realism through timestep-dependent guidance. VBench-Long evaluation results show that Diff-VF achieves a more favorable balance between temporal coherence and motion diversity than base models and recent training-free long video generation baselines, including FreeNoise, FreeLong, and RIFLEx, while maintaining competitive frame-wise quality. Experiments on two base models demonstrate the applicability to video diffusion models with different spatial-temporal modeling strategies. Extensive ablations validate the contribution of each component and hyperparameters.
Haoning Yang, Xinyuan Chen, Yaohui Wang +1
Shanghai Jiao Tong University, Shanghai, China · Shanghai Artificial Intelligence Laboratory, Shanghai, China
Extending the generation horizon of video diffusion models to long sequences remains a long-standing and important challenge. Existing training-free approaches fall into two categories: extensions of bidirectional models, which are tightly coupled to specific architectures and suffer from quality degradation over long horizons, and autoregressive models, which accumulate drift errors due to exposure bias and tend to produce repetitive motion patterns. To address these issues, we propose a novel but simple inference-time approach for long video generation that is architecture-agnostic and requires no additional training. Our method generates long videos via overlapping sliding windows, where predicted clean samples from adjacent windows are blended via \emph{Tweedie matching} to enforce both \textbf{manifold constraint and temporal consistency} across overlap regions. \emph{Stochastic early-phase sampling} then synchronizes per-window trajectories by injecting fresh noise after each Tweedie matching correction in the high-noise phase, before transitioning to deterministic ODE sampling to preserve fine-grained visual fidelity. Applied to various video generation models, our method generates videos several times longer than the native window length while outperforming both training-free and autoregressive baselines in temporal consistency and visual quality, and further extends to audio-video joint generation and text-to-3DGS without any fine-tuning.
Autoregressive (AR) video diffusion enables variable-length synthesis, but long-horizon generation often suffers from accumulated errors and identity drift. For efficiency, existing methods commonly adopt sliding-window attention during generation. This creates an irreversible generation trajectory: once the active window accumulates appearance errors, subsequent generations can only condition on this degraded trajectory and drift further away. We address this limitation by formulating long video generation as a retrieval-augmented generation (RAG) problem. Rather than relying solely on the recent window, we treat previously generated latents as a dynamic, searchable history. We propose LongLive-RAG, a general retrieval framework for AR video generation. At each new block, LongLive-RAG uses a query embedding to retrieve relevant historical latents. This lightweight retrieval step adds only a small overhead relative to generation and lets the generator condition on non-local context instead of only the recent window. To make retrieval more discriminative, we introduce the Window Temporal Delta Loss that suppresses redundant local similarity and encourages embeddings to capture meaningful temporal changes. Together, these components help reduce error accumulation caused by sliding-window attention. Experiments across multiple AR backbones and generation lengths show improved long-video quality and the best average VBench-Long rank. To our knowledge, among open-ended AR long video generation methods, LongLive-RAG is the first to formulate self-generated latent history as content-addressable retrieval memory. Code is available at https://github.com/qixinhu11/LongLive-RAG.