Autoregressive video models can generate minute-long videos in real time, but they produce generic subjects from text rather than specific subjects from user-provided images. Existing customization methods either require costly per-subject optimization or use pretrained conditioning networks that jointly process all video frames with bidirectional attention. Neither approach is designed for causal streaming. We present Custom Forcing, a training-free method that stores reference-based anchor frames in the persistent KV cache of a frozen autoregressive video model. However, fixed anchors face two limitations: simple conditioning allows identity to drift, and the text prompt continues to favor a generic subject. To address these problems, drift-adaptive value amplification (DVA) scales reference influence with the degree of identity drift, while anchor contrast guidance (ACG) steers generation away from the generic class prior. Over two-minute rollouts, fixed anchors fall from 0.58 to 0.42 in DINO-I, while Custom Forcing keeps it between 0.58 and 0.62 without reducing motion. Custom Forcing also achieves higher subject similarity than bidirectional customization methods and better preserves identity over 30s than causal image-to-video and reference-to-video models, while generating each frame 9.5-28.5 times faster than these long-video baselines.
Figures & tables
\fnum@figure : DINO-I over two minutes , n=100 . Custom K/V only uses fixed anchor frames without DVA or ACG.
\fnum@figure : Overview of Custom Forcing . Top: the reference images and copies of a custom anchor generated by FLUX.2 [klein] ( Black Forest Labs, 2025 ) are written into the KV cache once as the custom K/V. Middle: in every self-attention layer of the frozen backbone, the current chunk reads the cache with and without the anchor frames, giving z+ and z− , after DVA scales the reference values by g(t) . Bottom left: DVA sets g(t) from how far the similarity s(t) of each finalized chunk falls below its initial level μbase . Bottom right: ACG extrapolates away from z− by κ and blends the result with z+ .
DINO-I over time ↑
Alignment ↑
Video quality
Method
Length
Long video
Training free
all
first
middle
last
CLIP-I
CLIP-T
Subj. Cons. ↑
Bg. Cons. ↑
Dyn. Degree
Bidirectional
VACE-1.3B
5.1 s
✗
✓
0.333
–
–
–
0.703
0.296
0.921
0.951
0.46
SMRABooth
3.3 s
✗
✗
0.407
–
–
–
0.664
0.309
0.978
0.969
0.21
CustomCrafter
2.0 s
✗
✗
0.490
–
–
–
0.669
0.301
0.966
0.971
0.17
MotionBooth
3.0 s
✗
✗
0.435
–
–
–
0.642
0.293
0.907
0.928
0.76
Table 1: Main results. Long video marks whether a method can extend generation beyond a fixed length, and Training free whether it customizes the subject without training any model on it. Bold marks the best value in each column among the bidirectional methods and the 30 s rows, and the better of the two 2 min rows. Dynamic Degree is not bolded. Subject and Background Consistency favor the short bidirectional videos, which change little.
\fnum@figure : Two-minute generation. Custom K/V only drifts within the first minute, while Custom Forcing keeps both subjects for two minutes.
DINO-I over time ↑
Alignment ↑
Method
Backbone
Time
Causal
all
first
middle
last
CLIP-I
CLIP-T
FramePack
HunyuanVideo 13B
9.4×
✗
0.625
0.616
0.620
0.623
0.795
0.357
FramePack-F1
HunyuanVideo 13B
9.5×
✓
0.477
0.534
0.418
0.392
0.773
0.355
SkyReels-V2 DF
Wan2.1 1.3B
11.5×
✓
0.449
0.552
0.393
0.307
0.767
0.361
SkyReels-V3
Wan2.1 14B
28.5×
✓
0.641
0.693
0.620
0.564
0.794
0.347
Custom Forcing (Ours)
Wan2.1 1.3B
1.0×
✓
0.635
0.619
0.626
0.616
0.803
0.353
Table 2: Comparison with image- and reference-to-video models , 30 s, n=100 , identity and prompt alignment only. Backbone gives the architecture and size of each video diffusion transformer, Time is the generation time per frame relative to Custom Forcing on one H100 GPU, and Causal marks whether each part of the video is final before the next is generated. Bold marks the best score and underline the second best.
\fnum@figure : Qualitative comparison with image- and reference-to-video models. The image-to-video models start from the custom anchor, and SkyReels-V3 starts from the reference images, so its scenes differ. Custom Forcing keeps both subjects to the end.
DINO-I ↑
DVA
ACG
first
mid.
last
Custom K/V only
✗
✗
0.579
0.506
0.475
+ DVA
✓
✗
0.600
0.598
0.598
+ DVA + ACG (Ours)
✓
✓
0.619
0.626
0.616
Table 3: Component ablation , 30 s.
\fnum@figure : Adding the components one at a time. Red boxes mark where the subject departs from the reference images.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Value
Anchor slots
N
9
Reference anchors
∣R∣
5
Retained history frames
H
3
Chunk size / tokens per frame
F / L
3 / 1560
Injection budget
αinj
300
Drift cap
bmax
0.2
Appendix
Table 4: Hyperparameters of Custom Forcing . No value is tuned per subject or per prompt, and only b(t) adapts to each video.
DINO-I over time ↑
Alignment ↑
Video quality
Method
Length
DVA
ACG
all
first
middle
last
CLIP-I
CLIP-T
Subj. Cons. ↑
Bg. Cons. ↑
Dyn. Degree
Vanilla AR model
30 s
–
–
0.167
0.169
0.158
0.149
0.692
0.364
0.873
0.916
0.79
Custom K/V only
30 s
✗
✗
0.556
0.592
0.516
0.505
0.789
0.355
0.890
0.923
0.81
+ DVA
30 s
✓
✗
0.614
0.614
0.584
0.562
0.787
0.347
0.878
0.921
0.62
+ DVA + ACG (Ours)
30 s
✓
✓
0.641
0.634
0.623
0.615
0.798
0.348
0.889
0.923
0.87
Custom K/V only
2 min
✗
✗
0.518
0.586
0.466
0.447
0.780
0.354
0.859
0.911
0.84
Appendix
Table 5: Results on Deep Forcing , n=100 . Columns and bold follow Tab. 1 , and rows follow Tab. 3 .
\fnum@figure : Custom Forcing on Deep Forcing. The custom anchor is on the left of each row. (a) 30 s. Top: “A toy marching along a paved park path lined with autumn trees, …”. Bottom: “A backpack standing upright on a quiet city sidewalk lined with shopfronts, …”. (b) Two minutes. Top: “A toy rolling across a wide stone plaza in a historic square, …”. Bottom: “A dog walking along a quiet city sidewalk lined with shopfronts, …”. Custom Forcing keeps the red body, yellow belly, and large eyes of the toy and the red color and shape of the backpack to the end of the 30 s videos, and the white body, black visor, and blue circles of the robot and the red-brown coat and drooping ears of the dog for two minutes.
\fnum@figure : Effect of the ACG warm-up on Deep Forcing , Without warm-up, the generation remains close to the custom-anchor pose and exhibits limited motion. With warm-up, the subject shows greater motion while preserving its identity.
\fnum@figure : DINO-I over time , n=100 . Left: 30 s with each component. Right: 2 min.
DINO-I ↑
sprior↓
Method
first
middle
last
first
middle
last
Custom K/V only
0.579
0.506
0.475
0.395
0.414
0.418
+ ACG
0.595
0.551
0.550
0.380
0.385
0.378
Appendix
Table 8: Similarity to the reference images and to the class prior , 30 s, n=100 .
\fnum@figure : User study interface. Participants see the prompt, the reference images and choose A or B for each question before moving to the next pair.
Human
Qwen3-VL-235B
Claude Sonnet 5
Claude Opus 5.5
Custom Forcing vs.
Subject
Quality
Subject
Quality
Subject
Quality
Subject
Quality
Custom K/V only
60.0
45.8
60.0
50.0
40.0
35.0
80.0
60.0
FramePack
60.0
70.0
50.0
45.0
45.0
50.0
55.0
65.0
FramePack-F1
62.5
57.5
65.0
70.0
80.0
75.0
95.0
85.0
SkyReels-V2 DF
64.2
59.2
75.0
80.0
80.0
75.0
90.0
85.0
SkyReels-V3
45.8
51.7
80.0
90.0
70.0
80.0
60.0
80.0
Appendix
Table 9: User study and VLM judges. Percentage of judgments that prefer Custom Forcing over each method; 50% means no preference. Human: 120 votes per cell from 30 participants. VLM judges: Qwen3-VL-235B, Claude Sonnet 5, and Claude Opus 5.5, each with one judgment on each of 20 pairs per cell and 100 in All .
Table 19
\fnum@figure : Number of reference anchors at N=9 , 30 s, n=100 . Reference images are repeated in the shaded region. Left: Subject Consistency and Dynamic Degree, with the two axes aligned at the default ∣R∣=5 . Right: DINO-I.
DINO-I
all
first
middle
last
Dynamic
R5K4
Custom K/V only
0.541
0.579
0.506
0.475
0.41
Custom Forcing (Ours)
0.635
0.619
0.626
0.616
0.41
G5K4
Custom K/V only
0.571
0.601
0.550
0.523
0.34
Custom Forcing (Ours)
0.630
0.628
0.628
0.606
0.30
G8K1
Custom K/V only
0.556
0.577
0.538
0.512
0.40
Appendix
Table 12: What the custom K/V holds , 30 s, n=100 . R, G, and K count reference images, generated images, and copies of the custom anchor. Bold marks the best value in each column, and Dynamic Degree is not bolded.
DINO-I
CLIP-I
CLIP-T
Subj. Cons.
Bg. Cons.
Dynamic
Custom Forcing (Ours)
0.635
0.803
0.353
0.928
0.945
0.41
b≡0.05
0.651
0.804
0.335
0.887
0.926
0.28
b≡0.1
0.403
0.729
0.329
0.793
0.875
0.08
b≡0.2
0.158
0.643
0.309
0.691
0.823
0.02
no subject mask
0.622
0.803
0.352
0.913
0.936
0.30
Appendix
Table 13: Adaptive amplification and subject mask , 30 s, n=100 . The largest constant, 0.2 , equals the cap bmax . Bold marks the best value in each column, and Dynamic Degree is not bolded.
\fnum@figure : A constant amplification raises DINO-I but breaks the scene. Frames at 17.5 s of two prompts, whose DINO-I rises from 0.612 to 0.681 and from 0.482 to 0.530 .
Slug
Prompt
backpack (“backpack”)
00_park_bench
A cinematic video of a backpack resting on a weathered wooden park bench in an autumn park, orange leaves turning overhead and drifting past behind it.
01_forest_rock
A cinematic video of a backpack sitting on a flat mossy rock beside a winding forest trail, dappled sunlight flickering through the canopy and ferns swaying at the edges.
02_city_sidewalk
A cinematic video of a backpack standing upright on a quiet city sidewalk lined with shopfronts, traffic drifting past far behind it and shop awnings rippling in the breeze.
03_beach
A cinematic video of a backpack resting on flat open sand at a wide beach, fine sand streaming past it and waves breaking far behind.
04_mountain_ridge
A cinematic video of a backpack sitting on a rocky mountain ridge above a deep valley, wind tugging at its straps and clouds drifting across the valley below.
Appendix
Table 14: Evaluation prompts , grouped by subject. The noun used in the prompts is in parentheses.
Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework minimizes maximum mean discrepancy (MMD) in frozen self-supervised video representation spaces, using a hybrid Nyström--Monte Carlo estimator to balance approximation bias and sampling variance. Memory-efficient replay and gradient subsampling make this objective practical. Using the same architecture and initialization as Self-Forcing, our 1.3B model improves the VBench Total score from 83.80 to 84.64 while retaining 17 FPS. Removing auxiliary score models also enables 14B post-training on eight H200 GPUs. Beyond distillation, learning from reference videos enables the acquisition of new visual styles, semantic concepts, and spatial priors without a target-specific diffusion teacher.
Chi Zhang, Yueyi Liu, Shi Haoyang +5
College of AI, Tsinghua University · IAIR, Xi’an Jiaotong University · Xianghui Academy, Fudan University +1
Causal video generators must predict from the past, but they need not learn only from it. In streaming autoregressive video diffusion, each emitted segment becomes a commitment that future segments must preserve. Standard training, however, only asks each causal state to explain the present. This creates what we call a representation-level planning gap: states that fit the current segment may discard identity, layout, and motion information needed for a consistent future. We introduce Video-Mirai, a training-only method that closes this gap without changing causal inference: the generator rolls out causally, a frozen foresight encoder reads the completed rollout non-causally, and a lightweight predictor distills the resulting stopped-gradient targets into causal states. Future frames supervise representations, never generator inputs. At inference, the encoder and predictor are discarded, leaving the original architecture, per-step FLOPs, and KV-cache behavior unchanged. Video-Mirai improves a strong Causal-Forcing baseline on 5-second VBench from 83.8 to 84.6 in terms of Total Score. On 30-second rollouts beyond the training horizon, subject consistency improves from 84.9 to 88.5 and background consistency from 90.2 to 91.9. Ablations identify future-conditioned targets as the key ingredient, and probes show that future frames become more decodable from current features. Causality should constrain inference, not representation supervision. Our study highlights that visual autoregressive models need foresight. Project page: https://y0uroy.github.io/Video-Mirai.
Yonghao Yu, Lang Huang, Runyi Li +2
The University of Tokyo · National Institute of Informatics · Peking University
Autoregressive video diffusion models support real-time synthesis but suffer from error accumulation and context loss over long horizons. We discover that attention heads in AR video diffusion transformers serve functionally distinct roles as local heads for detail refinement, anchor heads for structural stabilization, and memory heads for long-range context aggregation, yet existing methods treat them uniformly, leading to suboptimal KV cache allocation. We propose Head Forcing, a training-free framework that assigns each head type a tailored KV cache strategy: local and anchor heads retain only essential tokens, while memory heads employ a hierarchical memory system with dynamic episodic updates for long-range consistency. A head-wise RoPE re-encoding scheme further ensures positional encodings remain within the pretrained range. Without additional training, Head Forcing extends generation from 5 seconds to minute-level duration, supports multi-prompt interactive synthesis, and consistently outperforms existing baselines. Project Page: https://jiahaotian-sjtu.github.io/headforcing.github.io/.
Jiahao Tian, Yiwei Wang, Gang Yu +1
AGI Lab, Westlake University · University of California at Merced · StepFun