WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents
Organizations: UC Santa Barbara · MIT CSAIL · MIT-IBM Watson AI Lab
Abstract
As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anomalies in these simulated environments, such as floating objects, traversable walls, or objects inconsistent with the surrounding scene. Multimodal AI systems, including vision-language models (VLMs) and vision-language-action models (VLAs), have shown potential for automating this task. However, 3D world auditing is complex, requiring the close coupling of two distinct capabilities: action, to navigate the 3D world and search for anomalies systematically and efficiently; and visual reasoning, to understand the environment and identify anomalies from multimodal observations. It remains largely unexplored whether multimodal agents can effectively couple these two capabilities, using visual reasoning to identify potential anomalies while taking actions to validate them. In this paper, we introduce WorldAuditBench, a benchmark for 3D world auditing comprising 213 anomaly tasks across 13 environments built with Unreal Engine 5 and Three.js, spanning five anomaly families. We evaluate five frontier models under a fixed exploration budget using two auditing paradigms: VLA-based exploration followed by VLM-based anomaly identification, and an end-to-end VLM agent in which visual reasoning directly guides action selection. Across the evaluated models and two paradigms, success rates range from 6.6% to 42.3%, substantially below human performance (83.4%). Through the task of world auditing, WorldAuditBench provides a testbed for studying how multimodal agents couple action and visual reasoning in interactive 3D environments, while highlighting current limitations in their ability to gather and interpret evidence during exploration.
Figures & tables
| Work | Task family | World interaction | Anomaly identification | Anomaly coverage |
| Offline visual benchmarks | ||||
| Video-MME ( 2025 ) | Question answering | ✗ | ✗ | — |
| GlitchBench ( 2024 ) | Game testing | ✗ | ✓ | Game glitches |
| VideoGameQA-Bench ( 2025 ) | Game testing | ✗ | ✓ | Game glitches |
| VideoGlitchBench ( 2026 ) | Game testing | ✗ | ✓ | Game glitches |
| Interactive environments and benchmarks | ||||
| Auditor | SP | IP | SpC | TC | SeC | Overall | Cost |
| Human | 89.8 | 81.3 | 76.3 | 78.1 | 88.6 | 83.4 | N/A |
| VLM auditors | |||||||
| GPT-6 Astra | 59.3 | 29.3 | 45.1 | 12.5 | 68.2 | 42.3 | 2.955 |
| Claude Opus 5 | 35.6 | 19.5 | 29.4 | 12.5 | 50.0 | 28.2 | 2.445 |
| Gemini 3.8 Flash | 49.2 | 22.0 | 29.4 | 12.5 | 50.0 | 32.4 | 1.382 |
| Muse Spark 1.3 | 22.0 | 12.2 | 13.7 | 5.0 | 22.7 | 15.0 | 0.791 |
| Setting | VLM | VLA–VLM | ||
| Default | 33.3 | 5.6 | ||
| (a) Starting distance | ||||
| Near ( ) | 38.1 | 7.1 | ||
| (b) Exploration budget | ||||
| 23.0 | 4.0 | |||
| 34.9 | 7.1 | |||
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Environment | Engine | # Tasks |
| Indoor | ||
| Residential House | Unreal Engine | 14 |
| Sponza Atrium | Three.js | 13 |
| Family House | Three.js | 15 |
| City | ||
| Utopian City | Unreal Engine | 15 |
| Record | Contents |
|---|---|
| Scene description | The scene’s setting and available interactions. |
| Anomaly category | One of the fifteen categories in the taxonomy. |
| Evaluation rubric | The target anomaly and its expected normal appearance or behavior. |
| Human validation | Quality judgment, difficulty rating, and comments on the inspected task. |
| Operation | Input | Effect |
| Navigation and Interaction | ||
| Move | Distance, direction | Move 0.3–4 m forward/backward; stop at obstacles. |
| Turn | Angle | Turn left or right by up to . |
| Look | Angle | Tilt the view up or down by up to . |
| Interact | None | Activate a reachable object under the crosshair. |
| Wait | Duration | Stay still for 0.5–5 s to observe changes over time. |
| Evaluation subset | Human-VLM agreement | |
|---|---|---|
| Agreement (%) | ||
| Overall | 91.72 | 0.645 |
| Unreal Engine | 92.86 | 0.644 |
| Three.js | 90.00 | 0.642 |
| Family | Observable inconsistency | Example |
|---|---|---|
| Static physics | Unsupported or floating object | p. C.5 |
| Intersection, overlap, or structural misalignment | p. C.5 | |
| Incorrect scale or proportions | p. C.5 | |
| Interactive physics | Traversal through an expected obstruction | p. C.5 |
| Unexpected obstruction of an admissible path | p. C.5 | |
| Anomalous motion following contact or interaction | p. C.5 |
| Auditor | Static physics | Interactive physics | Spatial consistency | Temporal consistency | Semantic consistency | Overall | Cost ($/task) |