PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
Organizations: University of Science and Technology of China · ByteDance · City University of Hong Kong
Abstract
Vision-Language Models (VLMs) have shown strong multimodal reasoning capabilities, yet whether they truly capture the physical consistency underlying real-world dynamics remains unclear. Existing benchmark paradigms often suffer from fragmented evaluation, focusing on isolated cognitive stages while overlooking the inherent synergy between perception, reasoning, and physical judgment. The lack of a holistic perspective limits the ability to diagnose whether VLMs can reliably evaluate the physical authenticity of emerging generative models. To address these issues, we introduce PhysVista, a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human seeing-reasoning-assessment process. PhysVista restores this loop by jointly evaluating physical state perception, physical dynamics reasoning, and physical plausibility assessment. It further distinguishes event-level reasoning and scale-level reasoning to enable fine-grained analysis of physical understanding. In addition, PhysVista incorporates both real-world and AI-generated videos, allowing evaluation across diverse domains and emerging generative scenarios. Extensive experiments across a diverse set of VLMs reveal substantial limitations in physical reasoning and plausibility assessment, highlighting a persistent gap between visual recognition and genuine physical understanding, and pointing toward more principled designs for physically grounded multimodal intelligence.
Figures & tables
| Model | Spatial | Camera | Scale | Uncertainty | Avg. Acc. | Localization | Score | Comparison |
|---|---|---|---|---|---|---|---|---|
| GPT-6 Sol OpenAI (2026b) | 81.00% | 74.00% | 35.00% | 60.00% | 68.59% | 53.03% | 0.6141 / 0.5586 | 54.00% |
| GPT-5.2 OpenAI (2026a) | 85.00% | 38.50% | 43.33% | 76.11% | 64.06% | 42.09% | 0.3328/0.3423 | 51.00% |
| Claude Opus 5.5 Anthropic (2026b) | 79.50% | 69.50% | 35.00% | 56.67% | 65.78% | 54.87% | 0.5702 / 0.5882 | 57.00% |
| Claude Opus 4.6 Anthropic (2026a) | 79.00% | 47.50% | 25.00% | 47.22% | 55.15% | 46.63% | 0.4482/0.4320 | 38.00% |
| Gemini 3.1 pro Google (2026) | 74.50% | 54.50% | 31.67% | 53.33% | 58.28% | 45.34% | 0.3087/0.3087 | 53.00% |
| Doubao Seed 2.0 Pro ByteDance (2026) | 77.50% | 59.50% | 30.00% | 50.00% | 59.69% | 47.48% | 0.4018/0.3949 | 57.00% |
| Model | Order | Mechanism | Violation | Prediction | Counterfactual | Critique | Avg. Acc. |
|---|---|---|---|---|---|---|---|
| GPT-6 Sol OpenAI (2026b) | 82.67% | 77.60% | 44.48% | 46.67% | 44.54% | 88.00% | 62.68% |
| GPT-5.2 OpenAI (2026a) | 41.33% | 56.77% | 33.22% | 48.15% | 33.61% | 71.00% | 49.86% |
| Claude Opus 5.5 Anthropic (2026b) | 90.67% | 77.08% | 51.60% | 58.52% | 49.58% | 87.00% | 67.89% |
| Claude Opus 4.6 Anthropic (2026a) | 52.00% | 68.75% | 41.47% | 48.15% | 36.97% | 82.00% | 57.61% |
| Gemini 3.1 pro Google (2026) | 66.67% | 74.48% | 40.94% | 40.00% | 38.66% | 76.00% | 60.20% |
| Doubao Seed 2.0 Pro ByteDance (2026) | 91.33% | 59.90% | 42.09% | 35.56% | 40.34% | 70.00% | 60.06% |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Camera Motion | Definition |
|---|---|
| Arc Left/Right | Horizontal circular movement around a subject, keeping it centered. |
| Pan Left/Right | Horizontal rotation of the camera on a fixed axis. |
| Tilt Up/Down | Vertical rotation of the camera on a fixed axis. |
| Translate Up/Down | Vertical displacement while adjusting pitch to maintain focus. |
| Static | The camera remains stationary with no movement or rotation. |
| Zoom Out | Decreasing the lens focal length to widen the field of view. |
| Spatial | Localization | Camera | Scale | Uncertainty |
|---|---|---|---|---|
| 200 | 200 | 200 | 60 | 180 |
| Order | Mechanism | Violation | Prediction | Counterfactual | Critique | Score | Comparison |
|---|---|---|---|---|---|---|---|
| 150 | 192 | 224 | 135 | 119 | 100 | 200 | 100 |
| Model | Exact-match Accuracy | Macro-F1 | Micro-F1 |
|---|---|---|---|
| GPT-5.6 Sol [ 32 ] | 20.54% | 46.60% | 44.48% |
| GPT-5.2 [ 31 ] | 14.73% | 35.93% | 33.22% |
| Claude Opus 5.5 [ 8 ] | 15.63% | 60.41% | 51.60% |
| Claude Opus 4.6 [ 7 ] | 12.05% | 48.41% | 41.47% |
| Gemini 3.1 pro [ 21 ] | 15.62% | 42.77% | 40.94% |
| Doubao Seed 2.0 Pro [ 15 ] | 14.29% | 49.12% | 42.09% |