PhysAlign: A Benchmark for Evidence-Grounded Role Alignment in Multimodal Physics Reasoning
Organizations: Sun Yat-sen University · Dalian University of Technology · Xi’an Jiaotong-Liverpool University
Abstract
A key challenge in physics diagram understanding is correctly associating visual information with the physical entities, relations, and conditions it describes. Even when a value, symbol, or other local element is accurately recognized, assigning it to the wrong entity or scope can distort the underlying physical premise and lead to incorrect reasoning. To systematically study this challenge, we introduce \textbf{PhysAlign}, a benchmark designed to assess whether multimodal models correctly associate information recognized from physics diagrams with its intended physical role. By disentangling visual recognition from physical-role assignment through localized probes and controlled variants, PhysAlign isolates correspondence errors from recognition failures. It contains 3,341 human-validated probes spanning 986 physics problems, enabling systematic evaluation of visual recognition and physical-role correspondence at scale. We further introduce five complementary evaluation metrics, including CAcc, GAcc, and JAcc, which provide a comprehensive assessment of models' ability to recognize diagram content, establish correct physical correspondences, and solve the underlying physics problem. Across our evaluated multimodal models, PhysAlign reveals a consistent gap between local visual recognition and physical-role grounding. Even when the queried content is correctly recognized, the conditional correspondence error rate remains 13.8% for GPT-6-Astra and rises to about 50.6% for InternVL3.5-8B. These findings indicate that strong perception alone does not ensure reliable physical interpretation, exposing a distinct grounding bottleneck that is largely hidden by answer-level accuracy and highlighting the need for future models to better align recognized visual evidence with its physical meaning.
Figures & tables
| (a) Source coverage | ||
| Source dataset | Parents | Parents [-2pt] with +GT |
| SeePhys ( Xiang et al., 2025 ) | 268 | 119 |
| LiveK12Bench ( Wang et al., 2026 ) | 250 | 88 |
| PhysElite ( Xu et al., 2026 ) | 220 | 79 |
| OlympiadBench a ( He et al., 2024 ) | 209 | 68 |
| Gaokao-MM-Physics ( Zong and Qiu, 2024 ) | 23 | 17 |
| Rule-based baselines ( ) | ||||
| Random candidates: 18.69% Nearest region: 76.82% | ||||
| Information controls and interface robustness | ||||
| Reference | No image | Reordered | ||
| Model | (pp) | |||
| Qwen3.5-4B ( Qwen Team, 2026 ) | 40.21 | 38.78 | 41.93 | |
| Qwen3.5-9B ( Qwen Team, 2026 ) | 40.52 | 43.53 | 43.02 | |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Recorded status or outcome | Probes |
| (a) Initial model-screening states | |
| Approved by model screening | 3,370 |
| Rejected by model screening | 1,128 |
| Further review required at initial screening | 1,089 |
| Incomplete / call error | 7 |
| Total generated candidates | 5,594 |
| (a) Sampling allocation | |||
| Audit source group a | Total probes | Audited | Coverage (%) |
| Local | 1,679 | 1,176 | 70.04 |
| PhysElite | 792 | 554 | 69.95 |
| OlympiadBench | 669 | 469 | 70.10 |
| Expansion-v3 | 123 | 123 | 100.00 |
| Phyx-OE | 40 | 40 | 100.00 |
| All parents | Paired-input coverage | ||||
| Domain group | Count | Share (%) | Paired | Unpaired | Rate (%) |
| Mechanics a | 478 | 48.48 | 208 | 270 | 43.51 |
| Electromagnetism a | 267 | 27.08 | 112 | 155 | 41.95 |
| Optics | 50 | 5.07 | 15 | 35 | 30.00 |
| Thermal physics / thermodynamics | 48 | 4.87 | 10 | 38 | 20.83 |
| Waves and acoustics | 35 | 3.55 | 9 | 26 | 25.71 |
| (a) Split-level density and paired-probe coverage | ||
| Split | Localized probes / parent | Paired-probe coverage (%) a |
| Development | 2.63 | 43.27 |
| Test | 3.45 | 14.78 |
| Full release | 3.39 | 16.55 |
| Model | Reasoning setting | Sampling / decoding | Output cap | Image |
| Main probe evaluation | ||||
| Qwen3.5-4B | Native thinking | Greedy | 8,192 | IMG02 |
| Qwen3.5-9B | Native thinking | Greedy | 8,192 | IMG02 |
| Qwen3.5-27B | Native thinking | Greedy | 8,192 | IMG02 |
| InternVL3.5-8B-HF | Explicit CoT | Greedy | 8,192 | IMG01 |
| Gemini 3.8 Flash | Provider default | 8,192 | IMG03 | |
| Model | ||||||
| Qwen3.5-4B | 40.21 | 53.97 | 32.09 | 46.51 | 59.46 | 40.54 |
| Qwen3.5-9B | 40.52 | 65.81 | 37.55 | 54.55 | 57.06 | 42.94 |
| Qwen3.5-27B | 84.11 | 84.14 | 68.12 | 80.29 | 80.97 | 19.03 |
| InternVL3.5-8B-HF | 23.52 | 2.12 | 1.05 | 33.36 | 49.37 | 50.63 |
| GPT-6 Astra | 81.94 | 89.44 | 77.13 | 83.31 | 86.23 | 13.77 |
| Gemini-3.8-flash | 90.71 | 85.20 | 70.59 | 79.33 | 82.85 | 17.15 |
| Model | Reading correct, grounding correct | Reading correct, grounding wrong | Reading wrong, grounding correct | Reading wrong, grounding wrong |
| Qwen3.5-4B | 32.09 | 21.88 | 14.42 | 31.62 |
| Qwen3.5-9B | 37.55 | 28.26 | 17.00 | 17.19 |
| Qwen3.5-27B | 68.12 | 16.01 | 12.17 | 3.70 |
| InternVL3.5-8B-HF | 1.05 | 1.07 | 32.32 | 65.57 |
| GPT-6 Astra | 77.13 | 12.32 | 6.19 | 4.37 |
| Gemini-3.8-flash | 70.59 | 14.61 | 8.74 | 6.05 |
| Model | Base (Raw) | +GT | (pp) |
| Qwen3.5-4B | 46.51 | 54.61 | +8.10 |
| Qwen3.5-9B | 54.55 | 56.53 | +1.98 |
| Qwen3.5-27B | 80.29 | 82.00 | +1.71 |
| InternVL3.5-8B-HF | 33.36 | 34.81 | +1.45 |
| GPT-6 Astra | 83.59 | 84.90 | +1.32 |
| Gemini-3.8-flash | 78.67 | 80.86 | +2.20 |
| Model | Repaired share | Harmed share | Repair rate | Harm rate | (pp) | |
| Qwen3.5-4B | 311/412 | 13.35 | 5.25 | 24.96 | 11.29 | +8.10 |
| Qwen3.5-9B | 311/412 | 8.30 | 6.32 | 18.26 | 11.59 | +1.98 |
| Qwen3.5-27B | 311/412 | 5.20 | 3.48 | 26.37 | 4.34 | +1.71 |
| InternVL3.5-8B | 311/412 | 11.17 | 9.73 | 16.77 | 29.16 | +1.45 |
| Baseline | GAcc | Scope and interpretation |
| Uniform random candidate | 18.69 | Chance level under the official aggregation weights |
| Nearest-region heuristic | 76.82 | Restricted to probes with a valid spatial distance |
| Model | Random | Nearest | Model | Model–Nearest (pp) |
| Qwen3.5-4B | 19.84 | 76.82 | 49.04 | |
| Qwen3.5-9B | 19.84 | 76.82 | 51.78 | |
| Qwen3.5-27B | 19.84 | 76.82 | 86.35 | +9.53 |
| InternVL3.5-8B | 19.84 | 76.82 | 30.30 |
| (a) Grounding on : 271 parents/394 probes | |||
| Model | Random | Model | Model–Random (pp) |
| Qwen3.5-4B | 18.76 | 38.03 | +19.27 |
| Qwen3.5-9B | 18.76 | 35.21 | +16.46 |
| Qwen3.5-27B | 18.76 | 73.22 | +54.47 |
| InternVL3.5-8B | 18.76 | 17.80 | |
| Model | Reference | No image | (pp) | Reordered | (pp) |
| Qwen3.5-4B | 40.21 | 38.78 | 41.93 | ||
| Qwen3.5-9B | 40.52 | 43.53 | 43.02 | ||
| Qwen3.5-27B | 84.11 | 63.72 | 84.48 | ||
| InternVL3.5-8B-HF | 23.52 | 49.87 | 7.24 | ||
| GPT-6 Astra | 81.94 | 75.57 | 84.27 | ||
| Gemini-3.8-flash | 90.71 | 76.90 | 90.82 |
| Model | (pp) | ||
| Qwen3.5-4B | 39.19 | 41.31 | |
| Qwen3.5-9B | 40.67 | 44.61 | |
| Qwen3.5-27B | 85.51 | 87.89 | |
| InternVL3.5-8B | 23.37 | 25.88 | |
| GPT-6 Astra | 83.05 | 85.53 | |
| Gemini-3.8-flash | 92.08 | 92.36 |
| Benchmark | General evaluation | Localized diagnosis | Utility | ||||
| Input Ctrl. | Interm. Eval. | Fixed Anchor | Rec./Grd. Split | Paired Rec. Ctrl. | Det. Local Score | Solve Link | |
| SeePhys [ Xiang et al., 2025 ] | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| SeePhys Pro [ Xiang et al., 2026 ] | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | |
| PhysicsArena [ Dai et al., 2025 ] | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ |
| MathVerse [ Zhang et al., 2024 ] | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| PhysAlign (Ours) | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |