OPIS: An Input-Grounded Benchmark for Multi-Object Memory in Video World Models
Authors: Hao Wang, Tao Yu, Liuzhou Zhang, HeXin Wang, Haopeng Jin, Yuxuan Zhou, Xinming Wang, Hongzhu Yi, +8 more
Organizations: USTC · Infrec · CASIA · HKUST · Tsinghua University · UCAS · The Chinese University of Hong Kong · Sun Yat-sen University · University of Waterloo · Fiveages · Jiangnan University · NUS · SUTD
Video world models must preserve the visual state of the world over time, but existing evaluation protocols often rely on generated histories, video reference, or selected revisit viewpoints that can confound the assessment of a model's true memory capability. To address this, we introduce OPIS, an input-grounded benchmark that strictly anchors the assessment to a fixed set of object instances from the initial observation for evaluating multi-object memory in video world models. The OPIS dataset comprises 500 cases across real-world, embodied-robotic, and game-world domains, providing dense object-level annotations for 12,672 rigid, articulated, and deformable instances. Our object-centric evaluator combines association and explicit visibility reasoning to hierarchically measure Object (O) Presence (P), Identity (I), and Structure (S), utilizing static or dynamic evaluation tracks based on object kinematics. Across eight image-to-video or camera-conditioned world models, our proposed OPIS scores range from 48.65 to 56.01. As the reference inventory grows from less than 20 to more than 40 objects, the Presence, Identity, and Structure scores show an overall decline, with the average Identity score falling from 40.22 to 23.11. The results demonstrate that preserving the particular object instances in the input is considerably harder than generating plausible visual elements.
Figures & tables
Figure 1: Comparison of evaluation paradigms and the input-grounded approach. Panels (a)–(d) illustrate evaluation using independent frames, generated prefixes, reference videos, and selected revisits. Panel (e) anchors both rollouts to the same object reference bank and geometry.
Method
Eval. protocol
Object- level
Fixed input set
Visib.- aware
Object struct.
VBench
F, PW
△
✗
✗
✗
VBench-2.0
F, PW
△
✗
△
△
MIND
F, VR, R
✗
✗
✗
✗
MBench
PW, R
△
✗
△
△
R2M-Bench
R, PW
△
✗
△
△
WorldTrace
PW, R
✗
✗
✗
✗
Table 1: Evaluation design comparison.
Figure 2: OPIS dataset composition and construction pipeline. Top: 500 cases span real-world, embodied-robotic, and game-world domains, with ten subcategories and rigid, articulated, and deformable objects. Bottom: candidate images undergo VLM-assisted selection, noun-phrase (NP) cleaning, instance annotation, geometry and task construction, and human review to produce a fixed input reference.
Figure 3: Input-grounded evaluation of object presence, identity, and structure. Generated observations are associated with the fixed input reference using appearance and geometric evidence. The examples distinguish a missing object ( C , when expected visible), an appearance change ( A to A′ ), and a structural change; valid viewpoint and motion changes are allowed.
Model
P ↑
I ↑
S ↑
CS (%)
OPIS ↑ [95% CI]
Image-to-video
Seedance 2.0
89.75
35.07
60.08
81.53
56.01 [52.13, 59.69]
H3-Max-Turbo
93.04
36.77
55.64
76.77
55.57 [52.20, 58.93]
Wan-3.0
88.95
34.47
50.58
75.86
51.81 [48.55, 55.03]
Gemini-Omni-Flash
93.49
30.03
49.30
70.94
50.43 [47.13, 53.76]
Camera-conditioned
Table 2: Formal OPIS results. P, I, S, and structural frame coverage CS are shown on a 0–100 scale. S already includes the coverage multiplier. All columns use the case–subcategory–domain hierarchy. OPIS reports a 95% stratified case-bootstrap interval.
Figure 4: Object-centric diagnostic views. Panel (a) reports OPIS across scene domains; each cell averages subcategories equally within a domain. Panel (b) shows the change in OPIS under aggregation variants relative to the strict score. Panels (c)–(e) show how reference-inventory size affects the memory components P , I , and S , structural evidence coverage CS , and the fraction of identity-valid frames containing at least one strict-Identity threshold failure, respectively. Density values use the reported common-case diagnostic.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Scenes
Instances
Static
Dynamic
Real-world scenes
200
8,082
3,565
4,517
Embodied / robotic
150
1,932
864
1,068
Game worlds
150
2,658
1,699
959
Total
500
12,672
6,128
6,544
Appendix
Table 3: Composition of OPIS. Static and dynamic columns indicate the evaluation track assigned to each reference instance.
Reference objects
P
I
S
Fail-I (%)
CS (%)
≤20
94.06
40.22
54.65
44.24
67.14
21–40
88.52
33.10
55.22
51.58
79.74
>40
86.63
23.11
49.42
67.55
91.60
Appendix
Table 4: Multi-object load across reference-inventory sizes. All score and rate columns use a 0–100 scale.
Model
Strict
No gate
No coverage
Minimum
Uniform
Gemini-Omni-Flash
50.43
55.48
55.75
46.72
57.61
H3-Max-Turbo
55.57
60.99
59.87
50.90
61.81
Wan-3.0
51.81
57.84
55.97
47.00
58.00
Seedance 2.0
56.01
61.08
59.62
50.93
61.63
SANA-WM
48.65
55.70
53.96
44.38
55.62
LingBot-World 2.0
50.36
55.78
59.97
45.60
56.92
Appendix
Table 5: Aggregation sensitivity on the reported evaluation set. No gate removes only the τS structural threshold; no coverage removes only CS ; minimum replaces the within-frame mean for both I and S after thresholding; uniform uses equal P/I/S weights. All perception and association evidence is held fixed.
Model
Real-world
Embodied
Game
Gemini-Omni-Flash
43.92
53.14
54.23
H3-Max-Turbo
52.13
53.92
60.65
Wan-3.0
49.31
52.19
53.93
Seedance 2.0
53.13
56.54
58.35
SANA-WM
44.62
49.21
52.11
LingBot-World 2.0
48.35
44.00
58.74
Appendix
Table 6: OPIS by scene domain. Each domain score averages its subcategories equally.
Model
Home
Natural
Public indoor
Urban
Gemini-Omni-Flash
36.89
54.23
32.57
51.99
H3-Max-Turbo
41.02
70.88
45.81
50.83
Wan-3.0
29.17
70.69
43.79
53.61
Seedance 2.0
35.33
69.12
47.73
60.34
SANA-WM
23.11
54.68
46.86
53.83
LingBot-World 2.0
27.55
66.33
40.68
58.84
Appendix
Table 7: Real-world subcategories: OPIS score.
Model
Industrial
Laboratory
Simulated
Gemini-Omni-Flash
43.50
65.41
50.50
H3-Max-Turbo
45.17
65.63
50.95
Wan-3.0
49.05
64.79
42.71
Seedance 2.0
51.73
67.21
50.68
SANA-WM
44.83
62.44
40.35
LingBot-World 2.0
40.70
48.72
42.57
Appendix
Table 8: Embodied subcategories: OPIS score.
Model
Cartoon
Pixel
Realistic
Gemini-Omni-Flash
54.68
55.03
52.99
H3-Max-Turbo
63.11
58.09
60.76
Wan-3.0
45.49
56.30
60.00
Seedance 2.0
54.24
52.10
68.71
SANA-WM
46.97
54.72
54.64
LingBot-World 2.0
56.06
66.83
53.32
Appendix
Table 9: Game subcategories: OPIS score.
Model
CP
CI
Static
Dynamic
Unknown
No-S
Gemini-Omni-Flash
97.39
97.32
17.35
67.33
24.27
3
H3-Max-Turbo
97.46
97.46
19.68
74.23
26.50
3
Wan-3.0
98.68
98.52
23.64
72.41
25.19
3
Seedance 2.0
97.20
97.20
18.99
79.19
28.59
2
SANA-WM
96.88
96.88
23.73
75.02
23.79
3
LingBot-World 2.0
91.88
91.88
24.59
65.23
35.62
1
Appendix
Table 10: Coverage audit. Percentage columns follow hierarchical averaging. No-S is a rollout count.
Figure 5: Temporal memory diagnostics. Each interval is scored independently using the same strict rules. Left: component scores, averaged equally across models after hierarchical aggregation. Right: structural frame coverage; gray curves show individual models and black shows their mean.
Scene subcategory
Number of cases
Home indoor
21
Natural outdoor
30
Public indoor
17
Urban outdoor
15
Total
83
Appendix
Table 11: AI-generated images in the real-world portion of OPIS.
Model
Resolution
FPS
Duration (s)
Stride
Samples
Wan-3.0
854×480
30
10.000
8
34
H3-Max-Turbo
864×496
24
10.042
8
28
Gemini-Omni-Flash
640×360
24
10.000
8
27
Seedance 2.0
832×480
24
10.125
8
28
SANA-WM
1280×704
16
10.000
8
18
LingBot-World 2.0
832×464
16
10.000
8
18
Appendix
Table 12: Recorded generation formats and evaluation sampling. The Samples column counts evaluated frames per rollout.
Check
N
Agreeing
Agreement (%)
Structural-claim repeatability
200
181
90.5
Association: exact assignment
150
139
92.7
Visibility: five labels
150
136
90.7
Identity: three labels
150
129
86.0
Structure: three labels
150
125
83.3
Appendix
Table 13: Evaluator-validity results. Repeatability uses 200 claim pairs. Each manual-audit dimension contains 150 decisions. Agreement is the number of agreeing judgments divided by N , expressed as a percentage.
Model
Subcategory
Released case
Visible change
P
I
S
OPIS
Seedance 2.0
Home indoor
home_indoor_0020
tabletop and lighting details drift
100.0
2.5
52.4
41.9
Gemini-Omni-Flash
Urban outdoor
urban_outdoor_0005
foreground bicycles and furniture drift
63.0
5.2
24.3
24.4
Wan 3.0
Home indoor
home_indoor_0006
cat, coffee table, and sofa area disappear
94.1
4.2
14.7
26.4
H3-Max-Turbo
Simulated environment
simulated_environment_0023
cabinetry and fixture layout changes
92.6
27.5
24.4
39.3
Echo-WM-Flash
Natural outdoor
natural_outdoor_0037
inflatable forms multiply and change shape
100.0
11.4
51.2
45.0
LingBot-World 2.0
Robot lab
robot_lab_setup_0029
tabletop objects and equipment are replaced
100.0
4.9
60.8
46.3
Appendix
Table 14: Eight visual case studies, one per model, spanning seven dataset subcategories. P, I, S, and OPIS are strict case-level scores on a 0–100 scale.
Figure 6: Eight visual case studies. Each tile pairs the reference input (left) with a frame from that model’s evaluated generation (right). The selected examples span seven subcategories and show visible instance-appearance or structural drift. Red boxes in the two archived panels indicate the evaluator observations used in the original comparison.
Figure 7: Strict case-level Presence (P), Identity (I), Structure (S), and OPIS scores for the eight displayed cases. The chart visualizes the same values reported in Table 14 .
Figure 8: Seedance 2.0 on home_indoor_0020: full-frame reference and generated-frame comparison.
Figure 9: Wan 3.0 on home_indoor_0006: full-frame reference and generated-frame comparison.
Figure 10: LingBot-World 2.0 on robot_lab_setup_0029: full-frame reference and generated-frame comparison.
Figure 11: Matrix-Game 3.5 on Realistic_style_3D_world_0001: full-frame reference and generated-frame comparison.