Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Meanwhile, directly reusing bidirectional score models in causal Distribution Matching Distillation (DMD) creates a mismatch between generation and scoring contexts. We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and context-aligned autoregressive DMD. CSF exploits contextual information asymmetry by varying the history while keeping the noisy target fixed: a noise-mixed-history student aligns its intermediate representations with those of a clean-history exponential-moving-average teacher. This self-supervised signal encourages the student to extract semantic information and improves cross-modal alignment. Context-aligned AR DMD shares the causal mask and prefix across generator sampling, fake-score training, and real-score evaluation to match generated and reference distributions under a block-conditional KL objective. With calibrated teacher guidance, it performs clean-prefix few-step distillation and then adapts to generated histories without switching objectives or requiring separate consistency distillation. At 480p, Salt++ improves visual and motion quality by 57% and 45% over OmniForcing on JavisBench under the same 4-step causal setting. A separate scale-wise post-training stage extends Salt++ to 4-step 1664×960 generation, outperforming bidirectional LTX-2 on six of seven reported metrics. Project page: https://xingtongge.github.io/Saltpp
Figures & tables
Figure 1: Comparison with recent causal post-training recipes. Prior recipes switch objectives between multiple stages and score the causal generator with bidirectional models (BI-DMD). Salt++ keeps one AR DMD objective throughout and only shifts its conditioning from clean context to generated rollout; a further scale-wise stage reaches 1664×960 with four generator calls.
Figure 2: Two roles of causal context in post-training. (a) CSF exploits information asymmetry in causal histories for self-supervised representation learning. (b) Mismatched (top) and aligned (bottom) contexts across the generator, real score, and fake score.
Figure 3: Score–context configurations in causal DMD. The generator always samples under a causal prefix, while the two scores may use bidirectional (BI) or causal (AR) visibility. Only AR–AR scores each block under the aligned context, estimating the gradient of Eq. 4 .
Figure 4: Qualitative comparison of 480p generation. Coast and diner examples compare LTX-2 Base (40 steps), OmniForcing, and both Salt++ routes (4 steps). The AR–AR DMD route retains fine scene and facial detail. Complete prompts and more comparisons are provided in Appendix B .
Model
Causal
Steps
VQ ↑
MQ ↑
AQ ↑
CLIP ↑
IB-AV ↑
Javis ↑
DeSync ↓
(a) 480p generation
LTX-2 Base ( HaCohen et al., 2026 )
No
40
1.884
0.566
4.986
0.311
0.239
0.200
0.608
AR teacher (CSF)
Yes
40
2.304
0.857
4.565
0.316
0.202
0.162
0.746
OmniForcing ( Su et al., 2026 )
Yes
4
1.807
0.699
4.718
0.303
0.163
0.124
0.745
Salt++ (TF-dCM route)
Yes
4
2.013
0.826
4.976
0.313
0.229
0.185
0.710
Salt++ (AR–AR DMD route)
Yes
4
2.838
1.010
4.991
0.316
0.184
0.146
0.759
Table 1: Main results on JavisBench-mini with official prompts. (a,b) JavisBench at 480p/960p; (c) VBench at 480p. At 480p, bold/underline indicate the best/second-best 4-step results. At 960p, bold marks the best result across all methods
Model
Causal
Steps
VQ ↑
MQ ↑
AQ ↑
IB-AV ↑
AVH ↑
Javis ↑
DeSync ↓
LTX-2 Base ( HaCohen et al., 2026 )
No
40
1.981
0.640
4.959
0.256
0.245
0.211
0.573
OmniForcing ( Su et al., 2026 )
Yes
4
1.774
0.607
4.613
0.158
0.152
0.121
0.734
Salt++ (TF-dCM route)
Yes
4
2.074
0.760
4.838
0.245
0.233
0.193
0.704
Salt++ (AR–AR DMD route)
Yes
4
2.820
0.999
4.997
0.192
0.186
0.151
0.696
Table 2: Quantitative results on the 1,000 LTX-2-enhanced JavisBench prompts at 480p.
Figure 5: Cross-modal alignment during AR teacher training. Causal Self-Flow versus clean teacher forcing over 19.2k iterations, evaluated with 40-step inference on the 1,000 JavisBench prompts. Panels (a–c) report audio–visual agreement and panel (d) reports audio–text agreement.
Figure 6: 960p generation: Fitness coach. Columns compare LTX-2 (40+3 steps), OmniForcing extrapolated to 960p, upsampled Salt++ 480p, and native Salt++ 960p. Full frames and detail windows show the finer facial detail of native 960p generation. More examples appear in Appendix B.3 .
Figure 7: Qualitative results of few-step initialization. Rows compare TF-dCM with BI–BI, BI–AR, and AR–AR TF-DMD. AR–AR retains significantly better scene structure and details.
(a) Objective
Scores
VQ ↑
MQ ↑
AQ ↑
CLIP ↑
IB-TV ↑
IB-TA ↑
DeSync ↓
TF-dCM
–
1.916
0.611
4.686
0.307
0.264
0.145
0.766
TF-DMD
BI/BI
0.837
0.140
4.796
0.301
0.266
0.144
0.811
TF-DMD
BI/AR
1.131
0.157
4.571
0.287
0.255
0.149
0.531
TF-DMD
AR/AR
3.147
1.351
4.750
0.317
0.271
0.153
0.726
Table 3: Few-step initialization ablations at 480p. (a) Score contexts and objectives on official prompts. (b) Teacher guidance under aligned AR–AR scores, on both prompt sets. Bold/underline mark best/second-best results in (a); bold marks the better result in (b).
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Blocks and contexts
K
Number of temporally aligned audio–video blocks
zk=(zkv,zka)
Video and audio latents of block k
m∈{v,a}
Modality index
y
Text condition
ck
Conditional information visible at block k
ck⋆
Clean teacher-forced context
Appendix
Table 4: Notation.
Figure 8: Additional 480p comparisons. (a) Bald eagle, (b) Rocky island, (c) Wellness host, and (d) Studio speaker. Each case uses the method order of Fig. 4 .
Figure 9: Additional 480p comparisons. (a) Race car, (b) Cat at a window, (c) Sunset rocks, and (d) Garden cottage. Method ordering follows Fig. 4 .
Figure 10: Training-time guidance calibration: podcast and arcade examples. The lower randomized setting preserves more tonal and texture detail than fixed high guidance.
Figure 11: Additional training-guidance comparisons: Microphone host and Studio speaker. The two AR–AR TF-DMD checkpoints use the same settings as Fig. 10 : fixed 4.0 versus independent U(1.0,3.5) teacher guidance at 3.2k, 4-step inference, and CFG =1 .
Figure 12: Additional 960p comparison: Wellness host. Full frames above and the corresponding facial and hair detail windows below, using the four-method column order of Fig. 6 .