As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anomalies in these simulated environments, such as floating objects, traversable walls, or objects inconsistent with the surrounding scene. Multimodal AI systems, including vision-language models (VLMs) and vision-language-action models (VLAs), have shown potential for automating this task. However, 3D world auditing is complex, requiring the close coupling of two distinct capabilities: action, to navigate the 3D world and search for anomalies systematically and efficiently; and visual reasoning, to understand the environment and identify anomalies from multimodal observations. It remains largely unexplored whether multimodal agents can effectively couple these two capabilities, using visual reasoning to identify potential anomalies while taking actions to validate them. In this paper, we introduce WorldAuditBench, a benchmark for 3D world auditing comprising 213 anomaly tasks across 13 environments built with Unreal Engine 5 and Three.js, spanning five anomaly families. We evaluate five frontier models under a fixed exploration budget using two auditing paradigms: VLA-based exploration followed by VLM-based anomaly identification, and an end-to-end VLM agent in which visual reasoning directly guides action selection. Across the evaluated models and two paradigms, success rates range from 6.6% to 42.3%, substantially below human performance (83.4%). Through the task of world auditing, WorldAuditBench provides a testbed for studying how multimodal agents couple action and visual reasoning in interactive 3D environments, while highlighting current limitations in their ability to gather and interpret evidence during exploration.
Figures & tables
Figure 1: Anomaly taxonomy of WorldAuditBench (Section 3.2 ). Representative cases illustrate five families of physical, spatial, temporal, and semantic inconsistencies.
Work
Task family
World interaction
Anomaly identification
Anomaly coverage
Offline visual benchmarks
Video-MME ( 2025 )
Question answering
✗
✗
—
GlitchBench ( 2024 )
Game testing
✗
✓
Game glitches
VideoGameQA-Bench ( 2025 )
Game testing
✗
✓
Game glitches
VideoGlitchBench ( 2026 )
Game testing
✗
✓
Game glitches
Interactive environments and benchmarks
Table 1: Comparison with related benchmarks and simulation environments. WorldAuditBench evaluates interactive anomaly identification across 15 categories grouped into five families.
Figure 2: Task construction in WorldAuditBench. We collect scenes, introduce anomalies, and iteratively refine tasks through human review of their quality.
Figure 3: Two auditing paradigms. (a) VLM agents reason and use tools during exploration. (b) VLA–VLM auditors analyze trajectories after VLA exploration. Both use the same judge.
Auditor
SP
IP
SpC
TC
SeC
Overall ↑
Cost ↓
Human
89.8
81.3
76.3
78.1
88.6
83.4
N/A
VLM auditors
GPT-6 Astra
59.3
29.3
45.1
12.5
68.2
42.3
2.955
Claude Opus 5
35.6
19.5
29.4
12.5
50.0
28.2
2.445
Gemini 3.8 Flash
49.2
22.0
29.4
12.5
50.0
32.4
1.382
Muse Spark 1.3
22.0
12.2
13.7
5.0
22.7
15.0
0.791
Table 2: Auditing performance and cost on WorldAuditBench. SR (%) and cost ($/task). SP: static physics; IP: interactive physics; SpC: spatial consistency; TC: temporal consistency; SeC: semantic consistency. Bold and underline mark the best and second-best results.
Figure 4: Auditing time per task.
Setting
VLM
VLA–VLM
Default
33.3
5.6
(a) Starting distance
Near ( 1/3 )
38.1
(+4.8)
7.1
(+1.6)
(b) Exploration budget
0.5×
23.0
(−10.3)
4.0
(−1.6)
1.5×
34.9
(+1.6)
7.1
(+1.6)
Table 3: Auditing factors. SR (%), Δ (pp).
Figure 5: Auditing multiple anomalies.
Figure 6: Anomaly exposure and success.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Environment
Engine
# Tasks
Indoor
Residential House
Unreal Engine
14
Sponza Atrium
Three.js
13
Family House
Three.js
15
City
Utopian City
Unreal Engine
15
Appendix
Table 4: Environments and task statistics in WorldAuditBench.
Record
Contents
Scene description
The scene’s setting and available interactions.
Anomaly category
One of the fifteen categories in the taxonomy.
Evaluation rubric
The target anomaly and its expected normal appearance or behavior.
Human validation
Quality judgment, difficulty rating, and comments on the inspected task.
Appendix
Table 5: Task annotations and human validation records.
Figure 7: Human validation interface. Reviewers inspect the environment alongside its scene description, anomaly category, and evaluation rubric, then record quality, difficulty, and comments.
Operation
Input
Effect
Navigation and Interaction
Move
Distance, direction
Move 0.3–4 m forward/backward; stop at obstacles.
Turn
Angle
Turn left or right by up to 180∘ .
Look
Angle
Tilt the view up or down by up to 90∘ .
Interact
None
Activate a reachable object under the crosshair.
Wait
Duration
Stay still for 0.5–5 s to observe changes over time.
Appendix
Table 6: Actions and tools available to the auditor. Movement, viewpoint changes, interaction, and waiting each cost one step; observation, memory, and reporting cost none.
Evaluation subset
Human-VLM agreement
Agreement (%) ↑
κ↑
Overall
91.72
0.645
Unreal Engine
92.86
0.644
Three.js
90.00
0.642
Appendix
Table 7: VLM judge agreement with human judgments. κ denotes Cohen’s kappa.
Family
Observable inconsistency
Example
Static physics
Unsupported or floating object
p. C.5
Intersection, overlap, or structural misalignment
p. C.5
Incorrect scale or proportions
p. C.5
Interactive physics
Traversal through an expected obstruction
p. C.5
Unexpected obstruction of an admissible path
p. C.5
Anomalous motion following contact or interaction
p. C.5
Appendix
Table 8: Anomaly taxonomy and corresponding in-context demonstrations.
Auditor
Static physics
Interactive physics
Spatial consistency
Temporal consistency
Semantic consistency
Overall
Cost ↓ ($/task)
Appendix
Table 9: Results by rendering engine. Success rate (%) and inference cost ($/task). Task counts are listed by anomaly family. Bold and underline mark the best and second-best results within each engine and approach. Pooled results are reported in Table 2 .
Nepal Applied Mathematics and Informatics Institute for research (NAAMII), Nepal · INSAIT, Sofia University, Bulgaria · State University of New York at Stony Brook, Korea