cs.CVOct 7, 2026

Real-Time Joint Audio-Video Generation by Parallel Adapter Composition

Authors: Jingyu Li, Xiaoxiao Xiang, Yiwen Guo

Organizations: LIGHTSPEED · Independent Researcher

Abstract

Deploying a joint audio-video diffusion transformer for real-time, interactive generation normally requires two essential modifications: block-autoregressive attention, so frames can be emitted before the whole clip is finished, and few-step sampling, so each block is cheap. Conventionally, the streaming video literature obtains both capabilities from a chained pipeline. It first distills a bidirectional teacher into a causal student, then into a few-step one, or proceeds in reverse order. Each stage of such a chain fine-tunes the weights the previous one produced, so a later objective can undo an earlier capability. Following the idea of model merging, we show that on a packed audio-video backbone the two capabilities can be acquired in parallel. A causal adapter is trained against the frozen backbone, and an off-the-shelf few-step adapter provides the few-step capability. As the two edit different functional axes, we predict, and then verify, that their weight-update directions are near-orthogonal, without any explicit orthogonality constraint during training. Orthogonal updates should combine without interfering, so parallel composition is a direct sum. The two adapters are simply added at inference, with no joint training, yielding few-step, streaming audio-video whose image quality tracks the bidirectional teacher. Compared to the chained baselines, the composed model matches or beats them on most metrics, making parallel composition a practical approach. The resulting streaming system generates joint audio-video in real time, ≈\approx26 fps at 480×832480\times832 without quantization, and sustains 30 s of continuous generation with stable image quality.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory

    Jul 29, 2026Yanbo Ding, Zhizhi Guo, Quanyue Song +4Audio-Video GenerationLong Video Generation

  2. Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation

    Apr 26, 2026Chunyu Li, Jiaye Li, Ruiqiao Mei +4Audio-Video GenerationVideo Diffusion Models

  3. DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

    Aug 31, 2026Jiashu Zhu, Yanhao Zheng, Ruitian Tian +7Audio-Video GenerationVideo Diffusion Models