WOVEN: Weaving Visual World Modeling into Multimodal LLMs
Organizations: Northwestern University · UNC Chapel Hill · Carnegie Mellon University · Amazon
Abstract
Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types. We first evaluate 38 frontier MLLMs (e.g., GPT-5.4 and Qwen3-VL-235B-A22B) and find a substantial and systematic deficit: even the strongest models fall far below humans, and the failures recur across model families and persist with scale. We then train MLLMs at multiple scales on WOVEN and find that they learn a shared capability that transfers broadly: training subsets of only about 2,000 items each collectively improve 22 of 26 external benchmarks by up to 27.3 percentage points, and WOVEN data can replace 30-50% of a task's own training data with comparable accuracy. Controlled comparisons further yield a training recipe for visual world modeling, validated prospectively on held-out benchmarks: select supervision by the reasoning operation it teaches rather than by the actions, scenes, or domains it shows, and prefer larger changes to the visual state for robustness. Our work establishes visual transition reasoning as a reusable foundation for systematic visual world-model training in MLLMs.
Figures & tables
| Family | Reasoning operation | Distribution | Reasoning operation | Distribution |
|---|---|---|---|---|
| Causal | Forward dynamics | Inverse dynamics | ||
| Counterfactual | Counterfactual removal | Counterfactual substitution | ||
| Physical | Outcome prediction | Cued prediction | ||
| Temporal | Temporal ordering | Temporal adjacency |
| Action type | Reasoning type | Perturbation | ||||||||||||||
| Overall | Exo | Perc | Insp | Navi | Mani | Fwd | Inv | Rmv | Sub | Otm | Cue | Ord | Adj | Geo | App | |
| Human | 92.3 | 92.8 | 92.8 | 92.6 | 90.6 | 92.8 | 92.8 | 92.8 | 91.1 | 92.2 | 93.3 | 91.1 | 92.8 | 91.7 | 97.3 | 98.7 |
| GPT-5.4 | 65.8 | 62.4 | 61.2 | 58.3 | 68.8 | 81.2 | 83.8 | 81.7 | 64.5 | 76.0 | 81.9 | 85.8 | 57.1 | 27.0 | 62.7 | 95.5 |
| Intern-S1-Pro | 63.2 | 56.2 | 52.8 | 53.7 | 70.1 | 87.1 | 77.6 | 77.5 | 66.7 | 73.3 | 61.5 | 69.1 | 63.4 | 26.5 | 33.3 | 79.2 |
| Qwen2.5-VL-72B | 58.9 | 53.7 | 47.7 | 47.8 | 66.4 | 83.4 | 74.7 | 74.7 | 54.8 | 72.0 | 59.4 | 76.0 | 49.4 | 28.0 | 38.5 | 83.0 |
| Gemma-4-31B | 57.8 | 50.8 | 49.4 | 53.1 | 61.2 | 74.3 | 76.3 | 76.6 | 63.6 | 64.5 | 64.9 | 70.5 | 30.8 | 25.3 | 59.0 | 95.5 |
| ID test | OOD scenes test | Perturbation test | |||||||||
| Training | Overall | Exo | Perc | Insp | Navi | Mani | Adj | Overall | Overall | Geo | App |
| Base | 26.4 | 23.8 | 27.0 | 23.5 | 30.0 | 28.8 | 25.7 | 27.8 | 21.1 | 22.6 | 19.7 |
| Full training set | 89.3 | 79.4 | 90.7 | 90.5 | 90.0 | 95.3 | 69.7 | 88.2 | 38.3 | 35.0 | 41.6 |
| Gain over the base model after training on one subset | |||||||||||
| perc_causal | +22.9 | +5.1 | +33.0 | +13.4 | +26.9 | +39.8 | +0.2 | +20.0 | +1.7 | +0.9 | +2.5 |
| perc_temporal | +6.6 | +8.4 | +10.9 | +4.9 | +8.0 | +1.3 | +29.9 | +2.9 | +4.1 | +5.6 | +2.7 |
| SAT | InPhyRe | MindCube | AEQA | TB-L | WM-AB | WPred | TOMATO | BSwan | ERQA | SpViz | CV-Bench | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Base | 37.3 | 22.0 | 37.2 | 43.4 | 16.6 | 31.0 | 34.8 | 31.6 | 61.8 | 33.8 | 27.7 | 69.3 |
| WOVEN SFT | 55.3 | 24.6 | 37.8 | 49.5 | 11.5 | 37.1 | 39.6 | 31.5 | 55.9 | 37.5 | 28.8 | 72.0 |
| WOVEN GRPO | 61.3 | 34.9 | 48.6 | 54.4 | 26.8 | 39.8 | 41.5 | 36.1 | 65.6 | 34.5 | 31.1 | 72.4 |
| +24.0 | +13.0 | +11.3 | +11.0 | +10.3 | +8.9 | +6.7 | +4.5 | +3.8 | +3.8 | +3.4 | +3.0 |
| Subset | World Prediction | Action EQA |
|---|---|---|
| perc_causal | ||
| insp_causal | ||
| navi_causal | ||
| mani_causal | ||
| insp_cf | ||
| mani_cf |
| Comparison mixture | Item | Questions | Base | Recipe | Comp. | CI | ||
|---|---|---|---|---|---|---|---|---|
| Temporal-heavy | (1) | Set A | 31.1 | 37.9 | 34.4 | 0.0002 | ||
| Uniform | (1) | Set A | 31.1 | 37.9 | 36.4 | 0.03 | ||
| Camera-heavy | (1) | Set C | 54.1 | 58.4 | 57.8 | 0.04 | ||
| Exogenous-heavy | (3) | Set A | 31.1 | 37.9 | 33.2 | 0.0002 | ||
| Exogenous-heavy | (3) | Set B | 70.3 | 71.0 | 71.5 | 0.04 |
Appendix figures & tables76 assets
Supplementary material from the paper’s appendix.
Appendix
| Measurement | Value |
| Mean absolute residual, far half of the image, edited vs. source | (median , p90 ) |
| Same measurement for the resize round-trip of the source frame (control) | |
| Ratio edited / control, per frame | median ; of below |
| Lag-1 spatial autocorrelation of the far-region residual | |
| Reduction of the far-region residual by the best global shift ( px, px steps, frames) | |
| Ratio edited / control by violation type ( types) | – , every type |
| Metric | Observed | Permutation null ( seeds) | |
|---|---|---|---|
| Patch-level balanced accuracy (%) | |||
| Image-level AUC |
| Metric | Observed | Permutation null ( seeds) | |
|---|---|---|---|
| Patch-level balanced accuracy (%) | |||
| Image-level AUC |
| Model | Family | Params |
|---|---|---|
| GPT-5.4 | GPT | N/A |
| GPT-5.2 | GPT | N/A |
| Qwen-VL-Max | Qwen-max | N/A |
| Qwen2.5-VL-3B | Qwen2.5 | 3B |
| Qwen2.5-VL-7B | Qwen2.5 | 7B |
| Qwen2.5-VL-32B | Qwen2.5 | 32B |
| Per-action type | Per-reasoning type | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Exo | Perc | Insp | Navi | Mani | Cue | Otm | Adj | Ord | Fwd | Inv | Rmv | Sub | Overall | |
| human | 92.8 | 92.8 | 92.6 | 90.6 | 92.8 | 91.1 | 93.3 | 91.7 | 92.8 | 92.8 | 92.8 | 91.1 | 92.2 | 92.3 |
| appearance | geometric | overall | |
| human | 98.7 | 97.3 | 98.0 |
| Annotator 1 | Annotator 2 | Pooled CI | |
|---|---|---|---|
| All questions | 94.0 | 93.6 | 93.8 |
| Main questions (200) | 94.0 | 93.5 | 93.8 |
| Perturbation questions (50) | 94.0 | 94.0 | 94.0 |
| Exogenous (43) | 95.3 | 93.0 | 94.2 |
| Inspective (43) | 95.3 | 93.0 | 94.2 |
| Manipulative (43) | 95.3 | 93.0 | 94.2 |
| Question | Annotator 1 | Annotator 2 | Pooled CI | Agreement |
|---|---|---|---|---|
| Scripted action carried out (yes) | 94 | 91 | 92.5 | 92% |
| Scripted action carried out (yes or partly) | 99 | 98 | 98.5 | — |
| Final state plausible | 99 | 98 | 98.5 | 97% |
| Video continuous | 98 | 97 | 97.5 | 97% |
| Follows the stated physical principle (exogenous, 20) | 100 | 100 | 100 | 100% |
| Annotator 1 | Annotator 2 | Pooled CI | |
|---|---|---|---|
| Edits carried out as instructed with nothing else changed | 149/150 | 147/150 | 98.7% |
| Questions with all three edits correct | 49/50 | 47/50 | 96.0% |
| ID | Claim | Test | Statistic (df) | Raw | Bonf. adj. | Effect size | |
| F.1 | Inspective below manipulative, frontier | 11 models | sign test | positive | gap range – pp | ||
| Mirror confusion, frontier | 7 families | 1-sided binomial vs. | each | each | |||
| Orientation concentration, GPT-5.4 | err | vs. uniform | ( ) | (med–large) | |||
| Orientation concentration, Qwen2.5-VL-7B | err | vs. uniform | ( ) | (small) | |||
| F.2 | Appearance vs. geometric scaling slope | 4 dense ladders, 11 models | OLS log-linear, per-ladder intercepts, slope | app , geo pts/dec ( , ; ) | , | , | slope difference pts/dec, CI |
| Within-model app–geo gap, Qwen2.5-VL-7B | two-proportion | pp |
| Set | QA count | Sources |
|---|---|---|
| ID reasoning type items | 6,228 | 594 |
| ID perturbation items | 1,116 | 320 |
| Scene-OOD reasoning type | 3,496 | 332 |
| Action type | #Comp | Reasoning type list |
|---|---|---|
| exogenous | 4 | Otm, Cue, Ord, Adj |
| perceptive | 4 | Fwd, Inv, Ord, Adj |
| inspective | 6 | Fwd, Inv, Rmv, Sub, Ord, Adj |
| navigative | 4 | Fwd, Inv, Ord, Adj |
| manipulative | 4 | Fwd, Inv, Rmv, Sub |
| Model | Family | Params | ID acc |
|---|---|---|---|
| GPT-5.4 | GPT | N/A | |
| Intern-S1-Pro | Intern | N/A | |
| Qwen2.5-VL-72B | Qwen2.5 | 72B | |
| Qwen3-VL-235B-A22B | Qwen3 | 235B-A22B (MoE) | |
| GPT-5.2 | GPT | N/A | |
| Gemma-4-31B | Gemma4 | 31B |
| Model | Fwd | Inv | Rmv | Sub | Otm | Cue | Ord | Adj |
|---|---|---|---|---|---|---|---|---|
| GPT-5.4 | 83.8 | 81.7 | 64.5 | 76.0 | 81.9 | 85.8 | 57.1 | 27.0 |
| GPT-5.2 | 70.7 | 78.6 | 53.2 | 56.3 | 78.8 | 85.4 | 46.4 | 27.1 |
| Qwen-VL-Max | 69.6 | 79.4 | 63.3 | 74.6 | 48.6 | 59.4 | 34.7 | 26.3 |
| Qwen2.5-VL-3B | 19.7 | 42.2 | 21.0 | 17.6 | 22.6 | 23.3 | 26.9 | 25.1 |
| Qwen2.5-VL-7B | 53.4 | 53.8 | 37.1 | 54.8 | 36.5 | 43.8 | 31.2 | 24.3 |
| Qwen2.5-VL-32B | 49.4 | 70.5 | 40.3 | 46.1 | 36.1 | 49.0 | 49.3 | 28.6 |
| Model | Exo | Perc | Insp | Navi | Mani |
|---|---|---|---|---|---|
| GPT-5.4 | 62.4 | 61.2 | 58.3 | 68.8 | 81.2 |
| GPT-5.2 | 58.9 | 52.3 | 50.7 | 63.4 | 67.4 |
| Qwen-VL-Max | 40.8 | 46.2 | 52.1 | 60.8 | 79.4 |
| Qwen2.5-VL-3B | 24.2 | 24.9 | 22.8 | 32.2 | 28.9 |
| Qwen2.5-VL-7B | 33.5 | 32.2 | 34.4 | 46.4 | 64.3 |
| Qwen2.5-VL-32B | 39.6 | 43.2 | 39.6 | 61.0 | 58.3 |
| Model | permanence | cohesion | solidity | gravity | support | inertia | collision | containment |
|---|---|---|---|---|---|---|---|---|
| GPT-5.4 | 56.2 | 61.8 | 69.4 | 75.0 | 62.5 | 60.4 | 43.8 | 70.1 |
| GPT-5.2 | 57.6 | 60.4 | 62.5 | 75.0 | 50.7 | 55.6 | 43.8 | 66.0 |
| Qwen-VL-Max | 29.9 | 44.4 | 38.9 | 56.9 | 41.7 | 47.2 | 18.8 | 48.6 |
| Qwen2.5-VL-7B | 37.5 | 37.5 | 28.5 | 47.2 | 34.0 | 41.0 | 21.5 | 20.8 |
| Qwen2.5-VL-32B | 41.0 | 42.3 | 36.4 | 44.2 | 37.6 | 41.6 | 32.4 | 41.8 |
| Qwen2.5-VL-72B | 54.2 | 51.4 | 45.8 | 76.4 | 58.3 | 59.0 | 32.6 | 52.1 |
| Model | h/ego | h/allo | hh/ego | hh/allo |
|---|---|---|---|---|
| GPT-5.4 | 87.2 | 80.9 | 82.6 | 74.3 |
| GPT-5.2 | 69.8 | 67.7 | 64.9 | 67.4 |
| Qwen-VL-Max | 85.1 | 79.2 | 81.2 | 72.2 |
| Qwen2.5-VL-7B | 68.8 | 65.6 | 66.7 | 56.2 |
| Qwen2.5-VL-32B | 66.7 | 59.0 | 60.8 | 46.9 |
| Qwen2.5-VL-72B | 91.3 | 84.4 | 85.8 | 72.2 |
| Model | overall | geo | app |
|---|---|---|---|
| GPT-5.4 | 79.1 | 62.7 | 95.5 |
| GPT-5.2 | 60.0 | 45.7 | 74.4 |
| Qwen-VL-Max | 48.9 | 33.7 | 64.2 |
| Qwen2.5-VL-3B | 21.1 | 22.6 | 19.7 |
| Qwen2.5-VL-7B | 37.0 | 23.5 | 50.5 |
| Qwen2.5-VL-32B | 37.4 | 30.1 | 44.6 |
| Model | manipulative | inspective | gap |
|---|---|---|---|
| GPT-5.4 | 81.2 | 58.3 | |
| GPT-5.2 | 67.4 | 50.7 | |
| Qwen-VL-Max | 79.4 | 52.1 | |
| Qwen2.5-VL-7B | 64.3 | 34.4 | |
| Qwen2.5-VL-32B | 58.3 | 39.6 | |
| Qwen2.5-VL-72B | 83.4 | 47.8 |
| Model | translation | rotation | gap |
|---|---|---|---|
| GPT-5.4 | 92.4 | 74.0 | |
| GPT-5.2 | 88.2 | 56.6 | |
| Qwen-VL-Max | 82.6 | 42.7 | |
| Qwen2.5-VL-7B | 56.6 | 17.0 | |
| Qwen2.5-VL-32B | 58.0 | 41.3 | |
| Qwen2.5-VL-72B | 77.1 | 42.7 |
| Model | Params | ID | overall | geo | app |
|---|---|---|---|---|---|
| Qwen2.5-VL-3B | 3B | 26.4 | 21.1 | 22.6 | 19.7 |
| Qwen2.5-VL-7B | 7B | 41.6 | 37.0 | 23.5 | 50.5 |
| Qwen2.5-VL-32B | 32B | 47.7 | 37.4 | 30.1 | 44.6 |
| Qwen2.5-VL-72B | 72B | 58.9 | 60.8 | 38.5 | 83.0 |
| Model | Params | ID | overall | geo | app |
|---|---|---|---|---|---|
| Qwen3-VL-8B-Instruct | 8B | 33.1 | 23.0 | 25.1 | 21.0 |
| Qwen3-VL-30B-A3B | 30B-A3B (MoE) | 42.8 | 34.1 | 25.6 | 42.5 |
| Qwen3-VL-32B-Instruct | 32B | 55.2 | 50.4 | 32.4 | 68.5 |
| Qwen3-VL-235B-A22B | 235B-A22B (MoE) | 58.5 | 55.5 | 32.8 | 78.1 |
| Model | Params | ID | overall | geo | app |
|---|---|---|---|---|---|
| Gemma-3-4B | 4B | 28.9 | 24.4 | 24.6 | 24.2 |
| Gemma-3-12B | 12B | 41.5 | 40.3 | 32.6 | 48.0 |
| Gemma-3-27B | 27B | 47.0 | 49.5 | 36.4 | 62.5 |
| Gemma-4-26B-A4B | 26B-A4B (MoE) | 43.9 | 34.1 | 25.4 | 42.8 |
| Gemma-4-31B | 31B | 57.8 | 77.2 | 59.0 | 95.5 |
| Model | Params | ID | overall | geo | app |
|---|---|---|---|---|---|
| Qwen3.5-9B | 9B | 30.1 | 20.6 | 20.8 | 20.4 |
| Qwen3.5-27B | 27B | 33.3 | 36.8 | 31.2 | 42.5 |
| Qwen3.5-35B-A3B | 35B-A3B (MoE) | 22.5 | 12.6 | 10.6 | 14.7 |
| Qwen3.5-122B-A10B | 122B-A10B (MoE) | 26.9 | 25.8 | 24.4 | 27.2 |
| Qwen3.5-397B-A17B | 397B-A17B (MoE) | 55.4 | 50.0 | 32.8 | 67.2 |
| Model | ID acc | ID rank | Geo acc | Geo rank | |
|---|---|---|---|---|---|
| ERNIE-4.5-VL-424B-A47B | 51.7 | 11 | 25.1 | 30 | |
| Qwen2.5-VL-7B | 41.6 | 20 | 23.5 | 33 | |
| Gemma-4-26B-A4B | 43.9 | 17 | 25.4 | 28 | |
| Llama-4-Scout | 32.6 | 25 | 17.0 | 35 | |
| Qwen3-VL-30B-A3B | 42.8 | 18 | 25.6 | 27 | |
| Seed-2.0-Mini | 54.3 | 10 | 27.1 | 19 |
| Model | Ord | Adj | Gap |
|---|---|---|---|
| GPT-5.4 | 57.1 | 27.0 | |
| GPT-5.2 | 46.4 | 27.1 | |
| Qwen-VL-Max | 34.7 | 26.3 | |
| Qwen2.5-VL-32B | 49.3 | 28.6 | |
| Qwen2.5-VL-72B | 49.4 | 28.0 | |
| Qwen3-VL-32B | 37.1 | 26.3 |
| Comparison | Dense acc | MoE acc | gap |
|---|---|---|---|
| Qwen3-VL-32B vs. 30B-A3B | 55.2 | 42.8 | |
| Gemma-4-31B vs. 26B-A4B | 57.8 | 43.9 | |
| Qwen3.5-27B vs. 35B-A3B | 33.3 | 22.5 |
| Pair | centered cos |
|---|---|
| Qwen2.5-VL-72B Qwen3-VL-235B-A22B | |
| GPT-5.4 GPT-5.2 | |
| Qwen3-VL-235B Qwen3-VL-32B | |
| ERNIE-424B GPT-5.4 | |
| Qwen3-VL-235B GPT-5.4 | |
| Qwen2.5-VL-72B GPT-5.4 |
| Perturbation probe | |||||
| Variant | ID | OOD | all | geo | app |
| Qwen2.5-VL-3B (base) | 26.4 | — | 21.1 | 22.6 | 19.7 |
| +full-mix SFT | 89.3 ( 62.9) | 88.2 | 38.3 ( 17.1) | 35.0 ( 12.4) | 41.6 ( 21.9) |
| +inverse-text SFT | 95.5 ( 54) | — | — | — | — |
| Qwen2.5-VL-7B (base) | 41.6 | — | 37.0 | 23.5 | 50.5 |
| +mv-balanced SFT | 50.1 ( 8.5) | 57.5 | 52.4 ( 15.4) | 38.4 ( 14.9) | 66.5 ( 16.0) |
| inspective | perceptive | navigative | manipulative | mean | |
|---|---|---|---|---|---|
| base 3B | |||||
| inverse-text SFT | |||||
| Source / Target | Perc Fwd | Insp Fwd | Navi Fwd | Mani Fwd |
|---|---|---|---|---|
| Base 3B (acc, %) | ||||
| perceptive ( ) | own | |||
| inspective ( ) | own | |||
| navigative ( ) | own | |||
| manipulative ( ) | own |
| Fwd | Sub | Rmv | Inv | Adj | Ord | |
|---|---|---|---|---|---|---|
| Base 3B (acc, %) | ||||||
| SFT | own |
| Fwd | Sub | Rmv | Inv | |
|---|---|---|---|---|
| Base 3B Insp (acc, %) | ||||
| Sub Insp | own | |||
| Base 3B Mani (acc, %) | ||||
| Sub Mani | own |
| orb_l (own) | orb_r (own) | arc_up (orth) | |
|---|---|---|---|
| base 3B (%) | |||
| 3B | |||
| 7B |
| rot_l | rot_r | |
|---|---|---|
| base 3B (%) | ||
| rot_l only (7B) | own | cross |
| rot_r only (7B) | cross | own |
| orb_l | orb_r | arc_up | |
|---|---|---|---|
| base 3B (%) | |||
| orb_l only (7B) | own | cross | cross |
| orb_r only (7B) | cross | own | cross |
| mv_fwd (own) | mv_bwd (own) | rot_l (cross) | rot_r (cross) | |
|---|---|---|---|---|
| base 3B (%) | ||||
| mv-balanced 7B (%) | ||||
| Variant | own_exo subset acc |
|---|---|
| (shuffled) | |
| (emergence-ordered) | |
| ( ) | pp |
| Normalization | same reasoning family | same action type | neither |
|---|---|---|---|
| raw (points) | |||
| relative | |||
| headroom | |||
| per-benchmark -score |
| SAT | DSI | MVP | WM-AB | TB-L | ViewSp | R2VLM | InPhyRe | ERQA | BSwan | CLEVRER | TVBench | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 7B SFT | — | — | — | — | — | — | ||||||
| 7B GRPO | — | — | ||||||||||
| 32B SFT | — | — | — | — | — | — |
| Model | 3DSRBench |
|---|---|
| Qwen3.5-4B base | |
| Orca-4B | ( ) |
| WOVEN (ours) | ( ) |
| Arm | Training data | Relation to the real data |
|---|---|---|
| Real | WOVEN perceptive causal dynamics (forward and inverse dynamics questions, four option orderings). | — |
| Per-source permutation | Same items, images, options, and answer-letter balance. Within each source, the four camera actions are mapped to the outcomes by a random permutation with no fixed point: the forward question for action has the outcome of as its answer, and the inverse question showing outcome has as its answer. The answer changes in of the examples. | Identical format; the action–outcome mapping differs across sources, so no consistent rule can be learned. |
| Global mirror | Same items, images, options, and answer-letter balance. One mapping for all sources: move forward move backward, rotate left rotate right . The answer changes in of the examples. | Identical format; a consistent, learnable rule with the direction of every action reversed. |
| A-OKVQA | four-option questions from the A-OKVQA training set ( Schwenk et al., 2022 ) (COCO images), each in its four option orderings, plus additional random orderings. | Multiple-choice training with a balanced answer-letter distribution, without WOVEN content. |
| Benchmark | Base | Real | Permutation | Mirror | A-OKVQA |
|---|---|---|---|---|---|
| SAT (circular) | 39.3 | 63.3 | 35.1 | 37.8 | 44.0 |
| SPAR-Bench view change (real images) | 26.4 | 54.9 | 27.5 | 35.8 | 21.8 |
| ActionEQA (circular) | 25.6 | 31.8 | 22.2 | 26.5 | 22.6 |
| WorldPrediction (circular) | 14.7 | 16.6 | 12.4 | 11.3 | 10.3 |
| WOVEN OOD | 36.2 | 45.9 | 36.8 | 32.9 | 32.6 |
| Benchmark | Real Permutation | Real Mirror | Real A-OKVQA |
|---|---|---|---|
| SAT (circular) | |||
| SPAR-Bench view change | |||
| ActionEQA (circular) | |||
| WorldPrediction (circular) | |||
| WOVEN OOD |
| Benchmark | Arm | s1 | s2 | s3 | Mean |
|---|---|---|---|---|---|
| SAT (circular) | Real | 65.3 | 63.3 | 61.3 | 63.3 |
| Per-source permutation | 35.3 | 36.7 | 33.3 | 35.1 | |
| Global mirror | 37.3 | 40.0 | 36.0 | 37.8 | |
| A-OKVQA | 42.7 | 44.7 | 44.7 | 44.0 | |
| SPAR-Bench view change | Real | 56.3 | 55.2 | 53.3 | 54.9 |
| Per-source permutation | 29.9 | 25.7 | 26.8 | 27.5 |
| Arm | Forward: real answer | Forward: mirrored answer | Inverse: real answer | Inverse: mirrored answer |
|---|---|---|---|---|
| Base | 24.7 | 27.1 | 25.3 | 28.8 |
| Real | 97.2 / 98.6 / 95.8 | 1.0 / 0.3 / 1.7 | 89.2 / 93.1 / 85.8 | 6.6 / 3.8 / 6.2 |
| Per-source permutation | 21.5 / 24.3 / 21.9 | 18.4 / 20.5 / 21.9 | 22.9 / 24.3 / 22.2 | 31.2 / 29.2 / 32.3 |
| Global mirror | 37.2 / 42.0 / 20.5 | 26.7 / 32.3 / 71.2 | 5.9 / 8.3 / 5.6 | 87.2 / 84.7 / 92.4 |
| A-OKVQA | 24.7 / 25.0 / 22.6 | 25.0 / 25.7 / 26.7 | 31.9 / 33.3 / 25.3 | 29.5 / 30.2 / 31.9 |
| Benchmark | Model | A | B | C | D |
|---|---|---|---|---|---|
| SAT | Correct answers | 50.0 | 50.0 | ||
| Base | 62.0 | 38.0 | |||
| Full mixture, SFT | 51.0 | 49.0 | |||
| Full mixture, SFT + GRPO | 49.0 | 51.0 | |||
| Perceptive causal dynamics, SFT | 52.3 | 47.7 | |||
| Perceptive causal dynamics, SFT + GRPO | 55.0 | 45.0 |
| Benchmark | Scoring | Base | Full mixture, SFT | Full mixture, SFT + GRPO | Perc. causal, SFT | Perc. causal, SFT + GRPO |
|---|---|---|---|---|---|---|
| SAT | Standard | 57.3 | 68.3 | 67.7 | 69.7 | 66.3 |
| Circular | 39.3 | 62.7 | 61.3 | 64.7 | 58.7 | |
| ActionEQA | Standard | 43.1 | 50.2 | 49.0 | 49.5 | 52.0 |
| Circular | 25.6 | 35.4 | 34.2 | 32.8 | 35.5 | |
| SPAR-Bench view change | Standard | 26.4 | 41.8 | 39.5 | 54.8 | 55.9 |
| Circular | 0.4 | 20.3 | 19.5 | 30.7 | 36.0 |
| Benchmark | Scoring | Full mixture, SFT | Full mixture, SFT + GRPO | Perc. causal, SFT | Perc. causal, SFT + GRPO |
|---|---|---|---|---|---|
| SAT | Standard | ||||
| Circular | |||||
| ActionEQA | Standard | ||||
| Circular | |||||
| SPAR view change | Standard | ||||
| Circular |
| Reasoning type | Base | Forward-only | CI | |
|---|---|---|---|---|
| Counterfactual substitution | 28.5 | 93.3 | ||
| Inverse dynamics | 43.3 | 60.7 | ||
| Forward dynamics (trained) | 37.8 | 87.4 |
| Reasoning type | Base | Forward-only | CI | |
|---|---|---|---|---|
| Counterfactual substitution | 24.8 | 78.1 | ||
| Counterfactual removal (answer: unchanged scene) | 27.8 | 55.2 | ||
| Inverse dynamics | 17.4 | 44.1 | ||
| Forward dynamics (trained) | 28.9 | 86.3 |
| Domain | Benchmark | Units | Scoring |
|---|---|---|---|
| Spatial | SAT | 150 | circular: a question counts only if both answer orders are correct |
| Spatial | ViewSpatial | 5,712 | question accuracy |
| Embodied | ActionEQA | 3,782 | question accuracy |
| Embodied | Robo2VLM | 6,676 | question accuracy |
| Physical | InPhyRe | 19,981 | question accuracy |
| Physical | CLEVRER | 12,000 | per-option accuracy (official stratified sample) |
| Benchmark | Base | Best candidate | Best gain CI | Null | ||
|---|---|---|---|---|---|---|
| SAT | 37.3 | Perc. causal, SFT | 0.002 | |||
| ViewSpatial | 35.3 | Mani. cf, SFT + GRPO | 0.002 | |||
| ActionEQA | 43.4 | Mani. causal, SFT | 0.002 | |||
| Robo2VLM | 33.5 | Mani. cf, SFT | 0.002 | |||
| InPhyRe | 22.0 | Insp. cf, SFT | 0.002 | |||
| CLEVRER | 58.2 | Insp. cf, SFT + GRPO | 0.002 |
| Benchmark | Subset | Seed 1 | Seed 2 | Seed 3 | Mean |
|---|---|---|---|---|---|
| SAT | Perc. causal | ||||
| ViewSpatial | Mani. cf | ||||
| ActionEQA | Mani. causal | ||||
| Robo2VLM | Mani. cf | ||||
| InPhyRe | Insp. cf | ||||
| CLEVRER | Insp. cf |
| Benchmark | Base | Seed 1 | Seed 2 | Seed 3 | Mean | CI | |
|---|---|---|---|---|---|---|---|
| WOVEN ID | 27.9 | 44.4 | 44.5 | 44.3 | 44.4 | ||
| WOVEN OOD | 33.0 | 45.2 | 44.3 | 43.9 | 44.5 | ||
| SPAR-Bench view change (real images) | 28.0 | 46.4 | 45.6 | 51.0 | 47.6 | ||
| SAT (circular) | 30.0 | 40.0 | 46.0 | 55.3 | 47.1 |
| Benchmark | Base | Mean of seeds | |
|---|---|---|---|
| WOVEN ID | 29.8 | 47.7 | |
| WOVEN OOD | 36.2 | 45.9 | |
| SPAR-Bench view change (real images) | 26.4 | 54.9 | |
| SAT (circular) | 39.3 | 63.3 |
| Model | Forward | Inverse | Joint | Indep. | Excess CI | Cycle | joint CI |
|---|---|---|---|---|---|---|---|
| Base | 21.3 | 40.4 | 8.0 | 8.1 | 19.0 | — | |
| Full mixture, SFT | 97.1 | 94.8 | 93.5 | 92.1 | 94.3 | ||
| Full mixture, SFT + GRPO | 96.6 | 96.9 | 94.1 | 93.7 | 94.4 | ||
| Perceptive causal, SFT | 75.3 | 76.6 | 65.5 | 61.8 | 71.3 | ||
| Perceptive causal, SFT + GRPO | 75.7 | 77.4 | 66.9 | 62.8 | 72.8 | ||
| Inspective causal, SFT | 73.7 | 75.8 | 61.9 | 59.3 | 69.0 |
| Subset | Recipe | Uniform | Temporal-heavy | Camera-heavy | Exogenous-heavy |
|---|---|---|---|---|---|
| Perceptive causal dynamics | 197 | 137 | 45 | 245 | 45 |
| Inspective causal dynamics | 197 | 137 | 45 | 245 | 45 |
| Navigative causal dynamics | 196 | 137 | 45 | 46 | 45 |
| Manipulative causal dynamics | 196 | 137 | 45 | 46 | 45 |
| Inspective counterfactual | 226 | 136 | 44 | 245 | 45 |
| Manipulative counterfactual | 226 | 136 | 44 | 46 | 45 |
| Source | Subset | Questions | Example |
|---|---|---|---|
| MMSI-Bench | Motion (Cam.) | 74 | “The images are taken continuously from a first-person perspective. In which direction are you moving?” |
| VLM4D (real videos) | object motion direction | 1,111 | “Which way does the person turn?” |
| VLM4D (real videos) | motion direction from a stated viewpoint | 260 | “From the camera perspective, which direction are the cyclists moving toward?” |
| PhysBench | relationships motion | 145 | “Moving in the direction indicated by the arrow in the picture, which of the following options are you most likely to see?” |
| PhysBench | dynamics manipulation | 83 | “Which of the following options presents the correct sequence to use the spoon to pick the objects in the green bowl and put them in the green pot?” |
| Source | Subset | Questions | What is asked |
|---|---|---|---|
| MVBench | counterfactual inference | 200 | which collision happens if an object is removed |
| MVBench | moving attribute | 200 | attributes of the objects that move or stay still |
| MVBench | moving count | 200 | how many objects of a kind are moving |
| MVBench | object existence | 198 | whether moving objects are still present at the end |
| IntPhys 2 | solidity, fixed camera | 104 | whether the event is physically possible |
| IntPhys 2 | immutability, moving camera | 136 | whether the event is physically possible |
| Source | Subset | Questions | What is asked |
|---|---|---|---|
| MMSI-Bench | Motion (Cam.) | 74 | how the camera moved or rotated across consecutive images |
| VLM4D | camera motion | 6 | which way the camera turned or moved in the video |
| SPAR-Bench | view change, ScanNet / ScanNet++ | 125 / 136 | which motion the camera made between two images |
| CameraBench | pan | 1,200 | whether the camera pans left or right |
| CameraBench | tilt | 703 | whether the camera tilts up or down |
| CameraBench | roll | 524 | whether the camera rolls clockwise or counterclockwise |
| Questions | Base | Recipe | Uniform | Temporal-heavy | Camera-heavy | Exogenous-heavy |
|---|---|---|---|---|---|---|
| Set D, PhysBench static (327) | 55.7 | 56.1 | 55.4 | 56.5 | 56.5 | 56.5 |
| Set D, MMSI-Bench static (559) | 28.4 | 29.0 | 28.6 | 29.0 | 29.0 | 29.6 |
| Set A (1,673) | 31.1 | 37.9 | 36.4 | 34.4 | 36.3 | 33.2 |
| Questions | Mixture | s1 | s2 | s3 | s4 | s5 | s6 | s7 | s8 | Mean |
|---|---|---|---|---|---|---|---|---|---|---|
| Set A | Recipe | 38.7 | 36.8 | 38.3 | 37.0 | 38.4 | 38.3 | 36.3 | 39.6 | 37.9 |
| Set A | Uniform | 38.4 | 36.9 | 35.2 | 34.5 | 38.0 | 36.1 | 36.2 | 36.1 | 36.4 |
| Set A | Temporal-heavy | 35.8 | 35.1 | 33.0 | 35.8 | 35.6 | 35.3 | 32.5 | 32.4 | 34.4 |
| Set A | Exogenous-heavy | 30.6 | 33.5 | 34.5 | 32.2 | 35.2 | 33.9 | 34.4 | 31.6 | 33.2 |
| Set B | Recipe | 70.3 | 71.1 | 71.5 | 70.7 | 70.6 | 71.0 | 71.3 | 71.1 | 71.0 |
| Set B | Exogenous-heavy | 71.4 | 71.2 | 71.2 | 70.7 | 71.8 | 72.6 | 71.5 | 71.6 | 71.5 |