SyncRA: Learning Temporal Correspondence in Omni-Modal Models
Organizations: University of Wisconsin–Madison · University of Alberta · Arizona State University
Abstract
Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible answers grounded in incorrect audio-visual pairings. We diagnose this problem through controlled temporal swaps, revealing that model answers do not reliably follow changes in these pairings. To address it, we propose Synchrony-Guided Representation Alignment (SyncRA), a lightweight method for strengthening temporal correspondence between audio and vision. Specifically, SyncRA contrasts intermediate audio-visual representations within each video, aligning matching moments while separating mismatched ones to capture local temporal correspondence within a shared global context. The objective derives supervision directly from existing input timing, requiring no additional annotations and leaving inference unchanged. We evaluate SyncRA across four open omni-modal models spanning different sizes and architectures on five public video benchmarks. SyncRA consistently outperforms answer-only fine-tuning across all model-benchmark combinations, while substantially improving the ability to track changing audio-visual pairings in controlled evaluations. These results demonstrate that lightweight, targeted supervision can effectively strengthen temporal correspondence and translate into broad improvements in audio-visual question answering.
Figures & tables
| Backbone | Arch. | Base | Vanilla SFT | SyncRA | |
|---|---|---|---|---|---|
| Qwen2.5-Omni | Dense | 2.25 | 11.00 | 20.75 | +9.75 |
| MiniCPM-o 4.5 | Dense | 35.50 | 57.25 | 62.25 | +5.00 |
| Qwen3-Omni | MoE | 13.25 | 34.00 | 62.50 | +28.50 |
| Nemotron | MoE | 0.50 | 6.00 | 16.50 | +10.50 |
| Benchmark | Base | Text CoT-SFT | Vanilla SFT | SyncRA | ||
| Qwen2.5-Omni-7B (dense, full SFT) | ||||||
| WorldSense | 3,172 | 45.27 | 47.92 | +2.11 | ||
| Daily-Omni | 1,197 | 61.74 | 63.16 | +3.43 | ||
| OmniVideoBench | 1,000 | 25.10 | 35.50 | +0.80 | ||
| LVOmniBench | 1,014 | 28.21 | 31.66 | +1.51 | ||
| AVUT-Human | 1,733 | 64.74 | 67.86 | +3.48 | ||
Appendix figures & tables36 assets
Supplementary material from the paper’s appendix.
Appendix
| Split | Questions | Source videos |
|---|---|---|
| Training | 9,500 | 1,367 |
| Validation | 500 | 415 |
| Total | 10,000 | 1,782 |
| Setting | Value |
|---|---|
| Optimizer | AdamW; learning rate ; ; ; weight decay 0.1 |
| Schedule | Linear decay; 80 warmup steps |
| Budget | 2 epochs; 792 optimizer steps per epoch; 1,584 steps total |
| Effective optimization batch | 12 examples: 2 processes microbatch 1 accumulation 6 |
| Precision and stabilization | bfloat16; float32 normalized contrastive features, similarities, and contrastive cross-entropy; gradient clipping 1.0; gradient checkpointing |
| Data order | Fixed example order; no additional shuffling |
| Backbone | Update scope |
|---|---|
| Qwen2.5-Omni-7B | Full understanding/Thinker fine-tuning, including text, visual, and audio modules; speech generation is not trained. |
| MiniCPM-o 4.5 | Full understanding fine-tuning; text-to-speech is disabled. |
| Qwen3-Omni-30B-A3B | Routed experts and router gates are frozen; other permitted modules are trained. Trainable backbone parameters: 2,715,593,328 / 31,719,205,488 (8.56%). |
| Nemotron-3-Nano-Omni-30B-A3B | Routed experts are frozen; routers, shared experts, and other permitted modules are trained. Trainable backbone parameters: 3,640,738,752 / 33,015,546,816 (11.03%). |
| Backbone | Identifier and revision |
|---|---|
| Qwen2.5-Omni | Qwen/Qwen2.5-Omni-7B ae9e1690543ffd5c0221dc27f79834d0294cba00 |
| MiniCPM-o 4.5 | openbmb/MiniCPM-o-4_5 503e754207c94da6bb26850b4469f367c9ea3582 |
| Qwen3-Omni | Qwen/Qwen3-Omni-30B-A3B-Instruct 26291f793822fb6be9555850f06dfe95f2d7e695 |
| Nemotron | nvidia/Nemotron-3-Nano-Omni-30B-A3B- Reasoning-BF16 e5e9932441de940c9a62185c870ea5bcd4cd24e2 |
| Backbone | Layer/depth | Training video and interval grid | ||
|---|---|---|---|---|
| Qwen2.5-Omni | 3,584 | 21/28 | 229,376 | 2 FPS, at most 256 frames; native 2 s intervals |
| MiniCPM-o 4.5 | 4,096 | 9/36 | 262,144 | 1 FPS, at most 180 one-second chunks; native 1 s intervals |
| Qwen3-Omni | 2,048 | 36/48 | 131,072 | 2 FPS, at most 256 frames; temporal patch size divided by actual sampling FPS |
| Nemotron | 2,688 | 26/52 | 172,032 | 2 FPS, at most 256 frames; timestamp-defined 2 s intervals |
| Backbone | Candidates | Fit/validate | Probe media scope | Selected |
|---|---|---|---|---|
| Qwen2.5-Omni | 7, 14, 21 | 200/415 | Full training clip; 2 FPS, at most 256 frames; native 2 s grid | 21 |
| MiniCPM-o 4.5 | 9, 18, 27 | 200/415 | First 60 s; native 1 s audio–video intervals | 9 |
| Qwen3-Omni | 12, 24, 36 | 198/408 | Full training clip; 2 FPS, at most 256 frames; actual temporal grid | 36 |
| Nemotron | 13, 26, 39 | 200/415 | First 120 s; 2 FPS; native 2 s intervals | 26 |
| Quarter depth | Half depth | Three-quarter depth | ||||
|---|---|---|---|---|---|---|
| Backbone | Block | R@1 | Block | R@1 | Block | R@1 |
| Qwen2.5-Omni | 7 | 11.44 | 14 | 16.47 | 21 | 20.00 |
| MiniCPM-o 4.5 | 9 | 78.61 | 18 | 74.39 | 27 | 36.44 |
| Qwen3-Omni | 12 | 13.07 | 24 | 28.20 | 36 | 28.30 |
| Nemotron | 13 | 13.30 | 26 | 14.06 | 39 | 12.91 |
| Backbone | Evaluation instruction | Mode/template setting |
|---|---|---|
| Qwen2.5-Omni | Direct option-letter answer | No additional thinking mode |
| MiniCPM-o 4.5 | Step-by-step reasoning; final line Final answer: X | enable_thinking=False ; use_tts_template=False |
| Qwen3-Omni | Step-by-step reasoning; final line Final answer: X | Instruct weights |
| Nemotron | Neutral question and options | enable_thinking=False |
| Benchmark | Evaluation focus | |
|---|---|---|
| WorldSense | 3,172 | Integrated real-world audio–video understanding |
| Daily-Omni | 1,197 | Everyday audio–video understanding and reasoning |
| OmniVideoBench | 1,000 | Multitask audio–video understanding |
| LVOmniBench | 1,014 | Long audio–video understanding |
| AVUT-Human | 1,733 | Audio-centered video understanding |
| Cue | O | A | V | AV |
|---|---|---|---|---|
| Cue 1 | ||||
| Cue 2 | ||||
| Cue 3 | ||||
| Cue 4 |
| Checkpoint | AllFour | Overall | O | A | V | AV |
|---|---|---|---|---|---|---|
| Qwen2.5-Omni (dense) | ||||||
| Base | 2.25 | 40.75 | 58.50 | 33.25 | 21.00 | 50.25 |
| Vanilla SFT | 11.00 | 54.19 | 73.75 | 51.50 | 29.75 | 61.75 |
| Text CoT-SFT | 6.25 | 46.75 | 62.00 | 43.00 | 30.25 | 51.75 |
| SyncRA | 20.75 | 64.31 | 77.75 | 62.75 | 46.25 | 70.50 |
| MiniCPM-o 4.5 (dense) | ||||||
| Backbone | Architecture | 95% interval | |
|---|---|---|---|
| Qwen2.5-Omni | Dense | ||
| MiniCPM-o 4.5 | Dense | ||
| Qwen3-Omni | MoE | ||
| Nemotron | MoE |
| AllFour | Overall accuracy | |||
|---|---|---|---|---|
| Comparator | 95% interval | 95% interval | ||
| Vanilla SFT | ||||
| Permuted pairs | ||||
| Fixed time codes | ||||
| Clip-level AV | ||||
| LVOmniBench | Daily-Omni | |||
|---|---|---|---|---|
| Checkpoint | AllFour | Overall | AllFour | Overall |
| Base | 7.89 | 50.00 | 14.51 | 57.87 |
| Text CoT-SFT | 5.26 | 43.75 | 13.58 | 56.48 |
| Vanilla SFT | 25.00 | 67.11 | 36.11 | 73.46 |
| Permuted pairs | 28.95 | 67.76 | 31.48 | 71.68 |
| Fixed time codes | 36.84 | 76.32 | 46.60 | 81.10 |
| R@1 (%) | MAE (s) | |||
|---|---|---|---|---|
| Space | Vanilla | SyncRA | Vanilla | SyncRA |
| Unprojected states | 13.39 | 96.91 | 7.176 | 0.044 |
| Trained SyncRA | 9.03 | 97.68 | 11.389 | 0.034 |
| Spectrum-matched random | 7.58 | 80.72 | 15.332 | 1.334 |
| Method | AllFour | Native R@1 | Daily-Omni |
|---|---|---|---|
| Vanilla SFT | 34.00 | 13.39 | 75.86 |
| Cosine-only | 40.50 | 17.92 | 78.20 |
| Pairwise logistic ranking | 46.25 | 27.47 | 78.36 |
| SyncRA (InfoNCE) | 62.50 | 96.91 | 78.95 |
| Checkpoint | Audio input | Unprojected states | SyncRA | complement |
|---|---|---|---|---|
| Vanilla | Original | 12.75 | 8.35 | 12.69 |
| Vanilla | Donor replacement | 9.28 | 6.10 | 9.10 |
| Vanilla | Two-interval delay | 7.82 | 6.48 | 7.76 |
| SyncRA | Original | 97.07 | 97.78 | 96.87 |
| SyncRA | Donor replacement | 96.74 | 97.41 | 96.59 |
| SyncRA | Two-interval delay | 96.45 | 97.47 | 96.21 |
| Vanilla | SyncRA | |||
|---|---|---|---|---|
| Space | Current interval | Content origin | Current interval | Content origin |
| Unprojected states | 7.82 | 7.86 | 96.45 | 0.01 |
| SyncRA | 6.48 | 6.15 | 97.47 | 0.02 |
| Fit evaluate | Base MAE | Vanilla MAE | SyncRA MAE | Difference | Paired 95% interval |
|---|---|---|---|---|---|
| Audio audio | 8.996 | 8.188 | 6.680 | ||
| Video video | 5.675 | 5.157 | 4.651 | ||
| Audio video | 13.133 | 12.685 | 15.741 | ||
| Video audio | 18.107 | 18.051 | 17.290 |
| Control | Base | Vanilla | SyncRA |
|---|---|---|---|
| Shuffled time labels: audio audio | 30.816 | 30.249 | 29.617 |
| Shuffled time labels: video video | 31.041 | 30.196 | 29.044 |
| Shuffled time labels: audio video | 39.298 | 54.523 | 30.236 |
| Shuffled time labels: video audio | 60.822 | 55.117 | 62.424 |
| Training-label-mean constant | 29.276 | 29.276 | 29.276 |
| Candidate pool | Space | Vanilla | SyncRA | Clip-level AV |
|---|---|---|---|---|
| All 100 sources | Unprojected states | 8.50 | 15.00 | 18.50 |
| All 100 sources | Clip-level AV | 5.50 | 36.50 | 19.00 |
| All 100 sources | SyncRA | 7.50 | 68.50 | 20.50 |
| 10 duration neighbors | Unprojected states | 34.50 | 41.50 | 50.00 |
| 10 duration neighbors | Clip-level AV | 30.00 | 65.00 | 54.00 |
| 10 duration neighbors | SyncRA | 31.50 | 86.00 | 57.50 |
| Supervision | Native R@1 | AllFour | Daily-Omni | |
|---|---|---|---|---|
| block | Block 24 | Block 36 | ||
| 24 | 68.37 | 39.87 | 60.00 | 77.36 |
| 36 (default) | 19.96 | 96.91 | 62.50 | 78.95 |
| Native R@1 | AllFour | Daily-Omni | |
|---|---|---|---|
| 0.03 | 90.85 | 58.00 | 78.53 |
| 0.1 (default) | 96.91 | 62.50 | 78.95 |
| 0.3 | 93.72 | 58.75 | 78.36 |
| Benchmark | Vanilla SFT | Permuted pairs | Fixed time codes | Clip-level AV | SyncRA |
|---|---|---|---|---|---|
| WorldSense | |||||
| Daily-Omni | |||||
| OmniVideoBench | |||||
| LVOmniBench | |||||
| AVUT-Human | |||||
| Avg-5 |
| Benchmark | Vanilla SFT | Permuted pairs | Fixed time codes | Clip-level AV | SyncRA |
|---|---|---|---|---|---|
| WorldSense | |||||
| Daily-Omni | |||||
| OmniVideoBench | |||||
| LVOmniBench | |||||
| AVUT-Human | |||||
| Avg-5 |
| Benchmark | Vanilla SFT | Permuted pairs | Fixed time codes | Clip-level AV | SyncRA |
|---|---|---|---|---|---|
| WorldSense | |||||
| Daily-Omni | |||||
| OmniVideoBench | |||||
| LVOmniBench | |||||
| AVUT-Human | |||||
| Avg-5 |
| Benchmark | Vanilla SFT | Permuted pairs | Fixed time codes | Clip-level AV | SyncRA |
|---|---|---|---|---|---|
| WorldSense | |||||
| Daily-Omni | |||||
| OmniVideoBench | |||||
| LVOmniBench | |||||
| AVUT-Human | |||||
| Avg-5 |
| Objective | Run 1 (%) | Run 2 (%) | Run 3 (%) | Mean SD |
| WorldSense ( ) | ||||
| Vanilla SFT | 1,595/3,172 (50.28) | 1,608/3,172 (50.69) | 1,615/3,172 (50.91) | |
| SyncRA | 1,670/3,172 (52.65) | 1,683/3,172 (53.06) | 1,666/3,172 (52.52) | |
| Permuted pairs | 1,593/3,172 (50.22) | 1,588/3,172 (50.06) | 1,602/3,172 (50.50) | |
| Fixed time codes | 1,583/3,172 (49.91) | 1,577/3,172 (49.72) | 1,572/3,172 (49.56) | |
| Clip-level AV | 1,595/3,172 (50.28) | 1,605/3,172 (50.60) | 1,598/3,172 (50.38) | |
| Objective | Run 1 (%) | Run 2 (%) | Run 3 (%) | Mean SD |
| WorldSense ( ) | ||||
| Vanilla SFT | 1,795/3,172 (56.59) | 1,807/3,172 (56.97) | 1,801/3,172 (56.78) | |
| SyncRA | 1,826/3,172 (57.57) | 1,824/3,172 (57.50) | 1,832/3,172 (57.76) | |
| Permuted pairs | 1,637/3,172 (51.61) | 1,628/3,172 (51.32) | 1,639/3,172 (51.67) | |
| Fixed time codes | 1,659/3,172 (52.30) | 1,664/3,172 (52.46) | 1,662/3,172 (52.40) | |
| Clip-level AV | 1,660/3,172 (52.33) | 1,657/3,172 (52.24) | 1,666/3,172 (52.52) | |
| Objective | Run 1 (%) | Run 2 (%) | Run 3 (%) | Mean SD |
| WorldSense ( ) | ||||
| Vanilla SFT | 1,777/3,172 (56.02) | 1,789/3,172 (56.40) | 1,788/3,172 (56.37) | |
| SyncRA | 1,815/3,172 (57.22) | 1,820/3,172 (57.38) | 1,823/3,172 (57.47) | |
| Permuted pairs | 1,820/3,172 (57.38) | 1,816/3,172 (57.25) | 1,815/3,172 (57.22) | |
| Fixed time codes | 1,806/3,172 (56.94) | 1,800/3,172 (56.75) | 1,807/3,172 (56.97) | |
| Clip-level AV | 1,810/3,172 (57.06) | 1,804/3,172 (56.87) | 1,805/3,172 (56.90) | |
| Objective | Run 1 (%) | Run 2 (%) | Run 3 (%) | Mean SD |
| WorldSense ( ) | ||||
| Vanilla SFT | 1,731/3,172 (54.57) | 1,720/3,172 (54.22) | 1,718/3,172 (54.16) | |
| SyncRA | 1,760/3,172 (55.49) | 1,773/3,172 (55.90) | 1,763/3,172 (55.58) | |
| Permuted pairs | 1,757/3,172 (55.39) | 1,760/3,172 (55.49) | 1,758/3,172 (55.42) | |
| Fixed time codes | 1,768/3,172 (55.74) | 1,765/3,172 (55.64) | 1,769/3,172 (55.77) | |
| Clip-level AV | 1,765/3,172 (55.64) | 1,752/3,172 (55.23) | 1,768/3,172 (55.74) | |
| Backbone | Benchmark | Base: (%) | Text CoT-SFT: (%) |
|---|---|---|---|
| Qwen2.5-Omni | WorldSense | 1,436/3,172 (45.27) | 1,520/3,172 (47.92) |
| Daily-Omni | 739/1,197 (61.74) | 756/1,197 (63.16) | |
| OmniVideoBench | 251/1,000 (25.10) | 355/1,000 (35.50) | |
| LVOmniBench | 286/1,014 (28.21) | 321/1,014 (31.66) | |
| AVUT-Human | 1,122/1,733 (64.74) | 1,176/1,733 (67.86) | |
| MiniCPM-o 4.5 | WorldSense | 1,750/3,172 (55.17) | 1,767/3,172 (55.71) |
| Backbone | Macro | Objective | Run 1 | Run 2 | Run 3 | Mean SD |
|---|---|---|---|---|---|---|
| Qwen2.5-Omni | Avg-5 | Vanilla SFT | 52.54 | 52.63 | 52.78 | |
| SyncRA | 55.05 | 54.87 | 54.83 | |||
| Permuted pairs | 53.20 | 53.07 | 53.38 | |||
| Fixed time codes | 52.57 | 52.41 | 52.43 | |||
| Clip-level AV | 53.16 | 53.31 | 53.08 | |||
| Avg-3 | Vanilla SFT | 62.31 | 62.40 | 62.74 |
| Backbone | Macro | Objective | Run 1 | Run 2 | Run 3 | Mean SD |
|---|---|---|---|---|---|---|
| Qwen3-Omni | Avg-5 | Vanilla SFT | 59.38 | 59.46 | 59.65 | |
| SyncRA | 61.34 | 61.48 | 61.47 | |||
| Permuted pairs | 60.92 | 60.80 | 60.82 | |||
| Fixed time codes | 60.39 | 60.25 | 60.53 | |||
| Clip-level AV | 60.97 | 60.85 | 60.90 | |||
| Avg-3 | Vanilla SFT | 70.06 | 70.15 | 70.30 |