Reasoning across videos requires aligning events, matching identities, comparing motion, and integrating partial observations. Evaluating these capabilities and testing how to improve them requires both reliable labels and targeted supervision. We introduce SYNCR, a simulator-grounded framework that connects these two needs through shared task generators. Built on Habitat, Kubric, and CLEVRER, SYNCR derives answers from environment state and provides 4,000 evaluation questions and 15,960 training questions over disjoint videos, spanning eight cross-video reasoning tasks. Visual ablations and human evaluation assess dependence on the supplied evidence and answer recoverability. Evaluation of 22 multimodal large language models reveals persistent difficulties in physical comparison and scene integration that increasing model size does not consistently resolve. Supervised fine-tuning raises Qwen3-VL-8B's average SYNCR accuracy from 32.6% to 61.6%, with gains extending to task configurations and video sources absent from training for those tasks. Transfer to real footage is most consistent for temporal ordering: accuracy improves by 9.0-20.5 percentage points on constructed Assembly101 and Panoptic ordering sets across three checkpoints spanning two model families and two model sizes, with additional gains on existing temporal reasoning benchmarks. These results establish SYNCR as a controlled setting for diagnosing cross-video reasoning failures, testing their learnability, and identifying where synthetic supervision transfers.
Figures & tables
Figure 1: The SYNCR framework. Eight tasks in four families: temporal alignment across streams, identity and geometry across viewpoints, physical and numerical comparison, and integration of partial scene observations. Simulator-derived answers support both evaluation and targeted training.
Model
Temp. Alignment
Spat. Tracking
Comp. Reasoning
Holi. Synthesis
Avg
Sync
Order
ReID
Meas
Num
Kin
Count
Route
Human
100.0
100.0
92.0
76.0
100.0
68.0
92.0
88.0
89.5
Gemini-3-Flash †
28.0
84.0
72.0
32.0
24.0
4.0
28.0
48.0
40.0
Gemini-3.1-Pro †
60.0
96.0
68.0
48.0
16.0
8.0
32.0
40.0
46.0
GPT-5.4 †
36.0
64.0
48.0
36.0
20.0
16.0
32.0
44.0
37.0
GPT-6 Astra †
88.0
100.0
88.0
52.0
60.0
20.0
48.0
60.0
64.5
Table 1: Zero-shot accuracy (%) on SYNCR. Avg is the unweighted mean across eight tasks. The 17 standard open-weight checkpoints and the Thinking checkpoint use 500 questions per task; proprietary models † and the four-person human majority vote use the same first 100. Bold and underline mark the best and second-best results among the 17 standard open-weight checkpoints. Chance is 25%, except for Num (20%).
Figure 3: Visual-evidence controls (zero-shot). For each item, the question and options stay fixed while the model receives all videos, one randomly selected video, or no video. Qwen3-VL-32B’s all-video and one-video results use the first 200 items, while its no-video result uses all 500; the latter is therefore not a paired comparison.
Figure 4: SYNCR accuracy after finetuning Qwen3-VL-8B. All bars use the same 200 held-out items per task. Each single-task SFT bar comes from a separate adapter trained on that task; eight-task SFT uses one adapter trained on all tasks. Eight-task RL applies GRPO to the base checkpoint, while SFT → RL applies it after eight-task SFT.
Figure 5: Real-video results after eight-task SYNCR SFT. Each cell shows the accuracy change, in percentage points, from that model’s own base on the same items. Blue indicates a gain and red a loss. Panel (a) uses SYNCR-format questions on real footage, while panel (b) shows individual tasks from existing benchmarks. Table 19 gives the absolute scores and item counts behind every cell.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Reported property
MVU- Eval
CV Bench
Cross Vid
MVP Bench
Gameplay QA
SYNCR (ours)
Evaluation QA pairs
1,824
1,000
9,015
5,050
2,365
4,000
Training QA pairs
0
0
0
0
0
15,960
Task categories/subtasks
8
15
10
14
15
8
Automatic QA construction
✓
✓
✓†
—
✓
✓
Simulator-state answer labels
—
—
—
—
—
✓
Direct no-video control
✓
—
—
—
✓
✓
Appendix
Table 3: Comparison with multi-video benchmarks. Columns refer to MVU-Eval ( Peng et al., 2025 ) , CVBench ( Zhu et al., 2025b ) , CrossVid ( Li et al., 2026 ) , MVPBench ( Bai et al., 2026 ) , and GameplayQA ( Wang et al., 2026 ) . Counts cover each full benchmark as defined by its authors; MVPBench and GameplayQA also include tasks that do not require multiple videos. In the training row, 0 means no separate training QA set is reported. A checkmark means the cited study reports the property; — means it does not report it. For automatic QA construction, a checkmark covers at least one task; † marks CrossVid’s eight automatically generated tasks out of ten. Controls differ in protocol.
Figure 6: Deterministic Ground Truth Generation in Habitat. Habitat provides synchronized RGB, depth, and semantic observations for each rendered camera pose, from which SYNCR derives visibility, instance labels, and navigation structure.
Evaluation
Training
Category
Task
Simulation Engine
Unique Videos
QA Pairs
Unique Videos
QA Pairs
Videos/Inp.
Clip Len. (s)
Sequential Ordering
CLEVRER
500
500
2,000
2,000
4
0.87
Temporal Alignment
Multi-Angle Synchronization
Kubric
1,500
500
4,119
2,000
3
3.0
Object Re-identification
Habitat
72
500
370
2,000
2
20.0
Spatial Tracking
Spatial Measurement
Kubric
558
500
2,385
2,000
2
5.0
Numerical Comparison
CLEVRER
1,000
500
2,969
2,000
2
5.12
Appendix
Table 4: SYNCR dataset statistics. Unique videos count distinct video files; Videos/Inp. is the number of videos per question; Clip Len. is the length of each clip shown to the model. Per-task counts are the distinct videos used by that task; the total is the union over tasks, which is smaller than their sum because some videos are shared between tasks. No video is shared between the evaluation and training splits. Total unique-video counts deduplicate files reused across tasks; the per-task counts sum to 5,688 for evaluation and 19,872 for training.
Figure 16: Tasks without consistent scaling gains within Qwen3-VL (4B, 8B, 32B).
Figure 17: Effect of reasoning-specialized post-training.
Temp. Alignment
Spat. Tracking
Comp. Reasoning
Holi. Synthesis
Trained on ↓ / Evaluated on →
Sync
Order
ReID
Meas
Num
Kin
Count
Route
Avg
Δ
Base (Qwen3-VL 8B)
31.5
55.0
43.5
27.0
27.0
21.0
30.5
25.5
32.6
—
Sync only
37.0 (+5.5)
49.0 ( − 6.0)
45.0 (+1.5)
28.5 (+1.5)
26.5 ( − 0.5)
21.0 (+0.0)
30.0 ( − 0.5)
24.5 ( − 1.0)
32.7
+0.1
Order only
26.0 ( − 5.5)
98.5 (+43.5)
45.0 (+1.5)
27.0 (+0.0)
26.5 ( − 0.5)
26.5 (+5.5)
28.5 ( − 2.0)
24.5 ( − 1.0)
37.8
+5.2
ReID only
23.0 ( − 8.5)
60.5 (+5.5)
73.0 (+29.5)
30.0 (+3.0)
26.5 ( − 0.5)
25.5 (+4.5)
33.0 (+2.5)
22.5 ( − 3.0)
36.8
+4.1
Meas only
28.5 ( − 3.0)
58.0 (+3.0)
45.0 (+1.5)
40.0 (+13.0)
21.5 ( − 5.5)
23.5 (+2.5)
28.0 ( − 2.5)
20.5 ( − 5.0)
33.1
+0.5
Appendix
Figure 18: Transfer matrix for single-task SFT. Accuracy (%) of Qwen3-VL-8B after training on the row task, evaluated on the column task (200 items each), with the change from the base in parentheses.
Qwen3-VL-8B
InternVL3.5-8B
Benchmark
Task
Base
8-task SFT
8-task RL
Base
8-task SFT
MVU-Eval
Knowledge-Intensive Reasoning
41.0
43.5 ( +2.5 )
40.0 ( −1.0 )
38.5
42.5 ( +4.0 )
MVU-Eval
Temporal Reasoning
68.0
75.5 ( +7.5∗ )
68.5 ( +0.5 )
60.0
76.5 ( +16.5∗ )
CrossVid
Behavioral Understanding
55.0
56.0 ( +1.0 )
56.0 ( +1.0 )
45.0
52.5 ( +7.5∗ )
CrossVid
Procedural Step Sequencing (MC)
74.5
84.5 ( +10.0∗ )
75.0 ( +0.5 )
50.5
78.5 ( +28.0∗ )
CVBench
Key-Action Recognition
63.6
69.3 ( +5.7 )
62.5 ( −1.1 )
61.4
67.0 ( +5.7 )
Appendix
Table 18: Sim-to-real transfer. Accuracy (%) on MVU-Eval ( Peng et al., 2025 ) , CrossVid ( Li et al., 2026 ) (200 items per task) and CVBench ( Zhu et al., 2025b ) (all items). Parentheses show changes from each model’s own base.
Qwen3-VL-8B
Qwen3-VL-4B
InternVL3.5-8B
Source
Task
n
base
SFT
Δ
base
SFT
Δ
base
SFT
Δ
Constructed on real footage
Assembly101
Order
200
35.5
44.5
+9.0
29.0
48.0
+19.0
34.0
48.0
+14.0
Panoptic
Order
200
28.5
47.5
+19.0
26.5
43.0
+16.5
32.0
52.5
+20.5
ReID
200
11.0
18.5
+7.5
17.0
22.5
+5.5
17.0
23.5
+6.5
EPFL
Order
200
33.0
42.5
+9.5
29.5
38.5
+9.0
30.5
45.0
+14.5
Appendix
Table 19: Absolute accuracy behind Fig. 5 . For every column of the figure: the number of evaluation items, base and 8-task SFT accuracy (%), and the change in percentage points. Both checkpoints of each model use the same items. For CVBench C. CR, Qwen uses 52 items and InternVL uses 55, as reflected in the reported accuracies. Changes are computed before rounding the displayed accuracies; paired uncertainty for the constructed ordering sets is reported in Table 20 .
Footage
Base
Model
Δ
95% CI
p
Qwen3-VL-8B, 8-task SFT
Assembly101
35.5
44.5
+9.0
[+2.5,+16.0]
0.015
Panoptic
28.5
47.5
+19.0
—
<0.003†
Qwen3-VL-8B, 8-task RL
Assembly101
35.5
38.5
+3.0
[−1.0,+7.0]
0.21
Panoptic
28.5
33.5
+5.0
—
—
Appendix
Table 20: Ordering on real footage. Accuracy (%) on 200 items per set. Chance is 25%. A dagger denotes a conservative upper bound on the exact paired p over all pairings consistent with the corrected marginal scores; dashes indicate statistics that require item-level predictions.
Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent video streams remains poorly understood. Existing multi-video benchmarks rely largely on human-annotated real-world footage, limiting the precision of spatial, temporal, and physical ground truth and making it difficult to diagnose model failures. We introduce SYNCR, a controlled synthetic benchmark for cross-video reasoning with programmatically verified grounding. Built using Habitat, Kubric, and CLEVRER simulator engines, SYNCR contains 4,000 multi-video question-answer pairs grounded in 4,827 unique videos. It evaluates MLLMs across eight tasks spanning four diagnostic pillars: Temporal Alignment, Spatial Tracking, Comparative Reasoning, and Holistic Synthesis. Our zero-shot evaluation of leading open- and closed-weight MLLMs reveals a substantial gap between current models and humans: the best model achieves only 64.5% average accuracy, compared to an 89.5% human baseline. Models perform relatively well on temporal ordering but struggle with precise physical and spatial reasoning, with the best model reaching only 29.8% accuracy on Kinematic Comparison. We further find that parameter scaling and reasoning-specialized post-training improve temporal alignment capabilities, but do not reliably address fine-grained physical tracking or global spatial synthesis. Finally, a sim-to-real correlation analysis suggests that SYNCR tracks model-level trends on a real-world multi-video benchmark.
Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy +1
Cross-Video Reasoning (CVR) has emerged as a critical frontier in multimodal intelligence, requiring models to retrieve, align, and aggregate evidence distributed across multiple videos. Current Multimodal Large Language Models (MLLMs) often struggle with CVR, as simple single-pass strategies encode multiple videos into a shared compressed context, potentially obscuring rare but critical evidence. In this paper, we propose AgentCVR, a multi-agent framework that treats CVR as an active evidence-acquisition task. AgentCVR employs a Master Agent to iteratively coordinate specialized Visual and Audio Agents for targeted evidence extraction. To ensure efficient training, we introduce Script-Simulated RL, which optimizes the agent's policy with LLM-generated semantic scripts and a lightweight text-based simulator, bypassing costly multimodal inference during online exploration. Experimental results on a comprehensive CVR benchmark show that AgentCVR outperforms single-pass baselines and achieves comparable performance to state-of-the-art closed-source systems, particularly in complex cross-video alignment and localization. To ensure reproducibility, our code is available at https://github.com/wang-jh24/AgentCVR.
Yilun Qiu, Jiahe Wang, Cilin Yan +4
Xiaohongshu Inc. · Tsinghua Shenzhen International Graduate School, Tsinghua University
Despite remarkable progress toward general-purpose video models, a critical question remains unanswered: how far are these models from achieving true multimodal reasoning? Existing benchmarks fail to address this question rigorously, as they remain constrained by straightforward task designs and fragmented evaluation metrics that neglect complex multimodal reasoning. To bridge this gap, we introduce CLVG-Bench, an evaluation framework designed to probe video models' zero-shot reasoning capabilities via Context Learning in Video Generation. CLVG-Bench comprises more than 1,000 high-quality, manually annotated metadata across 6 categories and 47 subcategories, covering complex scenarios including physical simulation, logical reasoning, and interactive contexts. To enable rigorous and scalable assessment, we further propose an Adaptive Video Evaluator (AVE) that aligns with human expert perception using minimal annotations, delivering interpretable textual feedback across diverse video context tasks. Extensive experiments reveal a striking answer to our central question: while state-of-the-art (SOTA) video models, such as Seedance 2.0, demonstrate competence on certain understanding and reasoning subtasks, they fall substantially short with logically grounded and interactive generation tasks (achieving success rates <25% and ~0%, respectively), exposing multimodal reasoning and physical grounding as critical bottlenecks. By systematically quantifying these limitations, the proposed method provides actionable feedbacks and a clear roadmap toward truly robust, general-purpose video models. CLVG-Bench and code are released here.