Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Meanwhile, directly reusing bidirectional score models in causal Distribution Matching Distillation (DMD) creates a mismatch between generation and scoring contexts. We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and context-aligned autoregressive DMD. CSF exploits contextual information asymmetry by varying the history while keeping the noisy target fixed: a noise-mixed-history student aligns its intermediate representations with those of a clean-history exponential-moving-average teacher. This self-supervised signal encourages the student to extract semantic information and improves cross-modal alignment. Context-aligned AR DMD shares the causal mask and prefix across generator sampling, fake-score training, and real-score evaluation to match generated and reference distributions under a block-conditional KL objective. With calibrated teacher guidance, it performs clean-prefix few-step distillation and then adapts to generated histories without switching objectives or requiring separate consistency distillation. At 480p, Salt++ improves visual and motion quality by 57% and 45% over OmniForcing on JavisBench under the same 4-step causal setting. A separate scale-wise post-training stage extends Salt++ to 4-step 1664×960 generation, outperforming bidirectional LTX-2 on six of seven reported metrics. Project page: https://xingtongge.github.io/Saltpp
Figures & tables
Figure 1: Comparison with recent causal post-training recipes. Prior recipes switch objectives between multiple stages and score the causal generator with bidirectional models (BI-DMD). Salt++ keeps one AR DMD objective throughout and only shifts its conditioning from clean context to generated rollout; a further scale-wise stage reaches 1664×960 with four generator calls.
Figure 2: Two roles of causal context in post-training. (a) CSF exploits information asymmetry in causal histories for self-supervised representation learning. (b) Mismatched (top) and aligned (bottom) contexts across the generator, real score, and fake score.
Figure 3: Score–context configurations in causal DMD. The generator always samples under a causal prefix, while the two scores may use bidirectional (BI) or causal (AR) visibility. Only AR–AR scores each block under the aligned context, estimating the gradient of Eq. 4 .
Figure 4: Qualitative comparison of 480p generation. Coast and diner examples compare LTX-2 Base (40 steps), OmniForcing, and both Salt++ routes (4 steps). The AR–AR DMD route retains fine scene and facial detail. Complete prompts and more comparisons are provided in Appendix B .
Model
Causal
Steps
VQ ↑
MQ ↑
AQ ↑
CLIP ↑
IB-AV ↑
Javis ↑
DeSync ↓
(a) 480p generation
LTX-2 Base ( HaCohen et al., 2026 )
No
40
1.884
0.566
4.986
0.311
0.239
0.200
0.608
AR teacher (CSF)
Yes
40
2.304
0.857
4.565
0.316
0.202
0.162
0.746
OmniForcing ( Su et al., 2026 )
Yes
4
1.807
0.699
4.718
0.303
0.163
0.124
0.745
Salt++ (TF-dCM route)
Yes
4
2.013
0.826
4.976
0.313
0.229
0.185
0.710
Salt++ (AR–AR DMD route)
Yes
4
2.838
1.010
4.991
0.316
0.184
0.146
0.759
Table 1: Main results on JavisBench-mini with official prompts. (a,b) JavisBench at 480p/960p; (c) VBench at 480p. At 480p, bold/underline indicate the best/second-best 4-step results. At 960p, bold marks the best result across all methods
Model
Causal
Steps
VQ ↑
MQ ↑
AQ ↑
IB-AV ↑
AVH ↑
Javis ↑
DeSync ↓
LTX-2 Base ( HaCohen et al., 2026 )
No
40
1.981
0.640
4.959
0.256
0.245
0.211
0.573
OmniForcing ( Su et al., 2026 )
Yes
4
1.774
0.607
4.613
0.158
0.152
0.121
0.734
Salt++ (TF-dCM route)
Yes
4
2.074
0.760
4.838
0.245
0.233
0.193
0.704
Salt++ (AR–AR DMD route)
Yes
4
2.820
0.999
4.997
0.192
0.186
0.151
0.696
Table 2: Quantitative results on the 1,000 LTX-2-enhanced JavisBench prompts at 480p.
Figure 5: Cross-modal alignment during AR teacher training. Causal Self-Flow versus clean teacher forcing over 19.2k iterations, evaluated with 40-step inference on the 1,000 JavisBench prompts. Panels (a–c) report audio–visual agreement and panel (d) reports audio–text agreement.
Figure 6: 960p generation: Fitness coach. Columns compare LTX-2 (40+3 steps), OmniForcing extrapolated to 960p, upsampled Salt++ 480p, and native Salt++ 960p. Full frames and detail windows show the finer facial detail of native 960p generation. More examples appear in Appendix B.3 .
Figure 7: Qualitative results of few-step initialization. Rows compare TF-dCM with BI–BI, BI–AR, and AR–AR TF-DMD. AR–AR retains significantly better scene structure and details.
(a) Objective
Scores
VQ ↑
MQ ↑
AQ ↑
CLIP ↑
IB-TV ↑
IB-TA ↑
DeSync ↓
TF-dCM
–
1.916
0.611
4.686
0.307
0.264
0.145
0.766
TF-DMD
BI/BI
0.837
0.140
4.796
0.301
0.266
0.144
0.811
TF-DMD
BI/AR
1.131
0.157
4.571
0.287
0.255
0.149
0.531
TF-DMD
AR/AR
3.147
1.351
4.750
0.317
0.271
0.153
0.726
Table 3: Few-step initialization ablations at 480p. (a) Score contexts and objectives on official prompts. (b) Teacher guidance under aligned AR–AR scores, on both prompt sets. Bold/underline mark best/second-best results in (a); bold marks the better result in (b).
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Blocks and contexts
K
Number of temporally aligned audio–video blocks
zk=(zkv,zka)
Video and audio latents of block k
m∈{v,a}
Modality index
y
Text condition
ck
Conditional information visible at block k
ck⋆
Clean teacher-forced context
Appendix
Table 4: Notation.
Figure 8: Additional 480p comparisons. (a) Bald eagle, (b) Rocky island, (c) Wellness host, and (d) Studio speaker. Each case uses the method order of Fig. 4 .
Figure 9: Additional 480p comparisons. (a) Race car, (b) Cat at a window, (c) Sunset rocks, and (d) Garden cottage. Method ordering follows Fig. 4 .
Figure 10: Training-time guidance calibration: podcast and arcade examples. The lower randomized setting preserves more tonal and texture detail than fixed high guidance.
Figure 11: Additional training-guidance comparisons: Microphone host and Studio speaker. The two AR–AR TF-DMD checkpoints use the same settings as Fig. 10 : fixed 4.0 versus independent U(1.0,3.5) teacher guidance at 3.2k, 4-step inference, and CFG =1 .
Figure 12: Additional 960p comparison: Wellness host. Full frames above and the corresponding facial and hair detail windows below, using the four-method column order of Fig. 6 .
Distilling video generation models to extremely low inference budgets (e.g., 2--4 NFEs) is crucial for real-time deployment, yet remains challenging. Trajectory-style consistency distillation often becomes conservative under complex video dynamics, yielding an over-smoothed appearance and weak motion. Distribution matching distillation (DMD) can recover sharp, mode-seeking samples, but its local training signals do not explicitly regularize how denoising updates compose across timesteps, making composed rollouts prone to drift. To overcome this challenge, we propose Self-Consistent Distribution Matching Distillation (SC-DMD), which explicitly regularizes the endpoint-consistent composition of consecutive denoising updates. For real-time autoregressive video generation, we further treat the KV cache as a quality parameterized condition and propose Cache-Distribution-Aware training. This training scheme applies SC-DMD over multi-step rollouts and introduces a cache-conditioned feature alignment objective that steers low-quality outputs toward high-quality references. Across extensive experiments on both non-autoregressive backbones (e.g., Wan~2.1) and autoregressive real-time paradigms (e.g., Self Forcing), our method, dubbed \textbf{Salt}, consistently improves low-NFE video generation quality while remaining compatible with diverse KV-cache memory mechanisms. Project page: https://xingtongge.github.io/Salt
Xingtong Ge, Yi Zhang, Yushi Huang +6
Hong Kong University of Science and Technology · Vivix Group Limited
Real-time interactive video generation requires low-latency, streaming, and controllable rollout. Existing autoregressive (AR) diffusion distillation methods have achieved strong results in the chunk-wise 4-step regime by distilling bidirectional base models into few-step AR students, but they remain limited by coarse response granularity and non-negligible sampling latency. In this paper, we study a more aggressive setting: frame-wise autoregression with only 1--2 sampling steps. In this regime, we identify the initialization of a few-step AR student as the key bottleneck: existing strategies are either target-misaligned, incapable of few-step generation, or too costly to scale. We propose \textbf{Causal Forcing++}, a principled and scalable pipeline that uses \emph{causal consistency distillation} (causal CD) for few-step AR initialization. The core idea is that causal CD learns the same AR-conditional flow map as causal ODE distillation, but obtains supervision from a single online teacher ODE step between adjacent timesteps, avoiding the need to precompute and store full PF-ODE trajectories. This makes the initialization both more efficient and easier to optimize. The resulting pipeline, \ours, surpasses the SOTA 4-step chunk-wise Causal Forcing under the \textit{\textbf{frame-wise 2-step setting}} by 0.1 in VBench Total, 0.3 in VBench Quality, and 0.335 in VisionReward, while reducing first-frame latency by 50% and Stage 2 training cost by \sim$$4\times. We further extend the pipeline to action-conditioned world model generation in the spirit of Genie3. Project Page: https://github.com/thu-ml/Causal-Forcing and https://github.com/shengshu-ai/minWM .
Min Zhao, Hongzhou Zhu, Kaiwen Zheng +6
Tsinghua University · ShengShu · Renmin University of China
In this work, we propose Mutual Forcing, a framework for fast autoregressive audio-video generation with long-horizon audio-video synchronization. Our approach addresses two key challenges: joint audio-video modeling and fast autoregressive generation. To ease joint audio-video optimization, we adopt a two-stage training strategy: we first train uni-modal generators and then couple them into a unified audio-video model for joint training on paired data. For streaming generation, we ask whether a native fast causal audio-video model can be trained directly, instead of following existing streaming distillation pipelines that typically train a bidirectional model first and then convert it into a causal generator through multiple distillation stages. Our answer is Mutual Forcing, which builds directly on native autoregressive model and integrates few-step and multi-step generation within a single weight-shared model, enabling self-distillation and improved training-inference consistency. The multi-step mode improves the few-step mode via self-distillation, while the few-step mode generates historical context during training to improve training-inference consistency; because the two modes share parameters, these two effects reinforce each other within a single model. Compared with prior approaches such as Self-Forcing, Mutual Forcing removes the need for an additional bidirectional teacher model, supports more flexible training sequence lengths, reduces training overhead, and allows the model to improve directly from real paired data rather than a fixed teacher. Experiments show that Mutual Forcing matches or surpasses strong baselines that require around 50 sampling steps while using only 4 to 8 steps, demonstrating substantial advantages in both efficiency and quality. The project page is available at https://mutualforcing.github.io.
Yupeng Zhou, Lianghua Huang, Zhifan Wu +7
1VCIP, School of Computer Science, Nankai University · 2Tongyi Lab · 2Tongyi Lab,3Peking University +1