Organizations: Shanghai AI Laboratory · The Hong Kong Polytechnic University · National University of Singapore · Nanyang Technological University · Zhejiang University
Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training offers a promising remedy, directly adapting it to joint audio-video generation is challenging. Heterogeneous multimodal rewards entangle learning signals and complicate credit assignment. Joint optimization of two modality towers is computationally expensive given their divergent dynamics. Moreover, synchronization evaluation difficulty depends on paired samples, preventing fair reward comparisons. We propose AV-GRPO, a modality-anchored online diffusion RL framework, and 5DAV, a decoupled, difficulty-controllable training dataset. AV-GRPO includes three key modules: (1) modality-anchored rollouts to disentangle learning signals and stabilize difficulty; (2) trajectory-locked frozen-tower optimization to reduce cost and reassign credit; (3) adaptive objectives and perturbation strengths tailored to modality-specific dynamics. This converts coupled multimodal preference learning into unimodal subproblems for precise reward attribution and better synchronization. Our 5DAV dataset decouples samples across five dimensions for systematic training. Experiments on JavisBench and VABench demonstrate AV-GRPO outperforms LTX-2.3 in generation quality, semantic alignment and cross-modal synchronization under LoRA and full fine-tuning. Ablations confirm our designs. Code and data: https://github.com/zhiyuxu03/AV-GRPO
Figures & tables
Figure 1: Overview of the AV-GRPO training pipeline.
Figure 2: Qualitative comparison of LTX-2.3 and AV-GRPO (full). For each example, rows 2–3 show the video frames and corresponding audio waveform generated by the original LTX-2.3 model; rows 4–5 show those generated by AV-GRPO (full). Bold text in the prompt highlights the visual events and sound cues that require audio–video synchronization.
AV-Quality
Text-Consistency
AV-Consistency
AV-Synchrony
Model
size
VQ ↑
AQ ↑
TV-IB ↑
TA-IB ↑
CLIP ↑
CLAP ↑
AV-IB ↑
AVHScore ↑
JavisScore ↑
DeSync ↓
AV-align ↑
-T2A+A2V
TempoTKn
1.3B
-
-
0.084
-
0.205
-
0.139
0.122
0.103
1.532
-
TPoS
1.0B
-
-
0.201
-
0.229
-
0.124
0.129
0.095
1.493
-
-T2V+V2A
ReWaS
0.6B
-
-
-
0.123
-
0.280
0.110
0.104
0.079
1.071
-
Table 1: Main results on JavisBench. Best results are in bold , second-best are underlined . VQ: Visual Quality, AQ: Audio Quality. ( ↑ : higher is better; ↓ : lower is better).
Model
Speech Q&N
Audio Aes
T-V Align
T-A Align
A-V Align
Lip Sync
DeSync ↓
Alignment
Express- siveness
Visual Realism
Audio Realism
Artistry
LTX-2.3
1.383
3.319
0.211
0.398
0.227
1.351
0.800
4.455
4.371
4.399
3.830
3.766
LTX-2.3+GDPO
1.465
3.452
0.210
0.424
0.243
1.439
0.726
4.481
4.409
4.413
3.849
3.784
LTX-2.3+AV-GRPO(lora)
1.487
3.553
0.217
0.446
0.261
1.585
0.594
4.512
4.450
4.402
3.865
3.792
LTX-2.3+AV-GRPO(full)
1.510
3.631
0.215
0.468
0.272
1.646
0.542
4.510
4.434
4.395
3.878
3.798
Table 2: Main results on VABench. Best results are in bold , second-best are underlined . For all metrics except DeSync, higher is better; for DeSync, lower is better.
Figure 3: Left : Radar chart comparing performance on the 5DAV and VGGSound datasets. Middle : Metrics are reported as relative percentage changes with respect to the baseline. Right : Comparison of model performance with different alternating step intervals.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 4: More qualitative results of AV-GRPO (full).
Recent advances in joint audio-video generation have been remarkable, yet real-world applications demand strong per-modality fidelity, cross-modal alignment, and fine-grained synchronization. Reinforcement Learning (RL) offers a promising paradigm, but its extension to multi-objective and multi-modal joint audio-video generation remains unexplored. Notably, our in-depth analysis first reveals that the primary obstacles to applying RL in this stem from: (i) multi-objective advantages inconsistency, where the advantages of multimodal outputs are not always consistent within a group; (ii) multi-modal gradients imbalance, where video-branch gradients leak into shallow audio layers responsible for intra-modal generation; (iii) uniform credit assignment, where fine-grained cross-modal alignment regions fail to get efficient exploration. These shortcomings suggest that vanilla RL fine-tuning strategy with a single global advantage often leads to suboptimal results. To address these challenges, we propose OmniNFT, a novel modality-aware online diffusion RL framework with three key innovations: (1) Modality-wise advantage routing, which routes independent per-reward advantages to their respective modality generation branches. (2) Layer-wise gradient surgery, which selectively detaches video-branch gradients on shallow audio layers while retaining those for cross-modal interaction layers. (3) Region-wise loss reweighting, which modulates policy optimization toward critical regions related to audio-video synchronization and fine-grained alignment. Extensive experiments on JavisBench and VBench with the LTX-2 backbone demonstrate that OmniNFT achieves comprehensive improvements in audio and video perceptual quality, cross-modal alignment, and audio-video synchronization.
Guohui Zhang, XiaoXiao Ma, Jie Huang +9
University of Science and Technology of China · JD Explore Academy · Peking University
Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attention, treating high-level semantics and low-level details in a fully entangled manner. This is suboptimal for talking head synthesis: while audio and facial motion are semantically correlated, their low-level realizations (acoustic signals and visual textures) follow distinct rendering processes. Enforcing joint modeling across all levels causes unnecessary entanglement and reduces efficiency. We propose Talker-T2AV, an autoregressive diffusion framework where high-level cross-modal modeling occurs in a shared backbone, while low-level refinement uses modality-specific decoders. A shared autoregressive language model jointly reasons over audio and video in a unified patch-level token space. Two lightweight diffusion transformer heads decode the hidden states into frame-level audio and video latents. Experiments on talking portrait benchmarks show Talker-T2AV outperforms dual-branch baselines in lip-sync accuracy, video quality, and audio quality, achieving stronger cross-modal consistency than cascaded pipelines.
Zhen Ye, Xu Tan, Aoxiong Yin +8
Hong Kong University of Science and Technology · Independent Researcher · Zhejiang University +3
Video is a rich representation of a physical event, capturing appearance, geometry, motion, and temporal evolution. Other modalities, such as 3D body motion or audio, encode narrower aspects of the same event. We find that joint multimodal diffusion transformers exhibit a corresponding asymmetry in cross-modal correspondence: companion modalities develop strong correspondences to video, but the reciprocal correspondences through which they constrain video remain substantially weaker. We express both directions as comparable correspondence distributions over video tokens and define their disagreement as the reciprocal correspondence gap. We introduce RecCAR, standing for Reciprocal Cross-modal Attention Regularization, a KL regularizer that uses the well-established video-to-modality correspondence as a fixed reference and aligns the weaker modality-to-video correspondence toward it. Across joint video-motion and video-audio generation, RecCAR improves the Human Anatomy score from 0.69 to 0.75 and reduces audio-video desynchronization from 0.804 to 0.752, while improving overall generation