Multimodal large language models (MLLMs) are increasingly envisioned as core components of scientific and engineering agents, yet their ability to interpret physical fields remains poorly understood. Existing physics benchmarks largely emphasize textbook problem solving or intuitive physical reasoning, leaving open whether MLLMs can infer physically meaningful information from continuous field observations. We introduce PhysFieldBench, a benchmark comprising 24 tasks and 1,160 evaluation examples across controlled equation fields, simulated physical fields, and observed physical fields. The tasks assess three forms of inference: identifying physical mechanisms, comparing latent control variables, and predicting outcome properties. Across representative open-source and proprietary MLLMs, zero-shot performance is low: the best model achieves a chance-normalized score of 29.3, while several open-source models remain near chance. In contrast, a task-specific supervised vision transformer performs substantially better, demonstrating that the inputs contain learnable physical information. To diagnose these failures, a structured self-explanation analysis attributes most errors to missed visual patterns and incorrect visual-to-physical mappings. Further, to explore whether post-training can improve physical inference and generalize to unseen tasks, we compare supervised fine-tuning with final answers or chain-of-thought supervision and reinforcement learning. Final-answer supervision performs best overall but transfers less effectively, whereas reinforcement learning after chain-of-thought supervision achieves the best generalization. Together, these findings highlight the need to improve visual-to-physical grounding and cross-task generalization for MLLMs to reliably interpret physical fields in scientific and engineering workflows.
Figures & tables
Figure 1: While existing physics benchmarks focus primarily on textbook problems or intuitive reasoning, PhysFieldBench targets the physical-field understanding required by scientific and engineering agents: inferring mechanisms, control variables, and outcomes from diverse physical fields.
Figure 2: Overview of PhysFieldBench. It comprises 24 tasks and 1,160 examples from nine sources, spanning controlled, simulated, and observed physical fields and three inference types: mechanism, control, and outcome. Asterisks denote the six held-out leaf tasks used to assess cross-task transfer.
Figure 3: Construction pipeline of PhysFieldBench. Heterogeneous physical-field data first undergo physical-label verification, then are converted into a unified visual question-answering format with field-appropriate representations and further refined by physics-based filtering and human review.
Models
Field Domains
Inference Axes
Overall
Rank
Controlled
Simulated
Observed
Mechanism
Control
Outcome
Supervised Visual Reference
ViT ( 2021 )
54.8
76.3
80.1
63.6
73.3
69.5
70.4
-
Open-source Models
Qwen3-VL-8B-Instruct ( 2025 )
-1.3
-4.3
13.6
6.0
-3.0
9.5
2.7
11
LLaVA-OneVision-2-8B ( 2026 )
0.1
-7.4
1.5
2.2
-1.2
-5.0
-1.9
12
Table 1: Main results on PhysFieldBench. We report chance-normalized scores, macro-averaged over leaf tasks within each field domain and inference axis. Overall reports the macro-average over all tasks. A score of 0 indicates chance-level performance. MLLMs are evaluated zero-shot, while ViT is trained separately for each task on i.i.d. training data drawn from the same distribution as the test set. Bold and underlined values denote the best and second-best scores among MLLMs, respectively.
Figure 4: Self-explanation showcases. Gemini 3 Flash correctly integrates geometric and multi-view evidence in ShapeNetCar but shows physical-mapping and factor-disentanglement errors in 2D PDE.
Figure 6: Post-training results. (a) Normed score on training-covered and held-out tasks. (b) Raw Pass@1 accuracy and Pass@8 coverage (%) after reasoning supervision and reinforcement learning.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
ID
Task name
Source
Axis
Split
n
K
r
Controlled equation fields
C01
Single Dominant Mechanism Identification
PDEFormer1D
M
T
50
5
1/5
C02
Dominant Mechanism Pair Identification
PDEFormer1D
M
H
60
6
1/6
C03
Diffusion Coefficient Comparison
PDEFormer1D
C
T
40
2
1/2
C04
Nonlinear Advection Coefficient Comparison
PDEFormer1D
C
T
40
8
1/8
C05
Single Dominant Mechanism Identification
PDEFormer2D
M
T
40
4
1/4
Appendix
Table 2: Complete task inventory of PhysFieldBench. M, C, and O denote mechanism, control, and outcome inference, respectively. T and H denote training-covered and held-out tasks. n is the number of evaluation questions, K is the number of answer choices, and r=1/K is the uniform random baseline. Task IDs are used throughout the appendix.
Setting
Answer-only SFT
CoT SFT
CoT + GRPO
Checkpoint
Final, epoch 5
Step 185
Step 320
Decoding
Greedy
Greedy
Greedy
Thinking
Disabled
Enabled
Enabled
Maximum image pixels
589,824
589,824
589,824
Maximum generated tokens
64
9,216
9,216
Truncated responses
0
8
4
Appendix
Table 6: Inference settings for the selected post-trained Qwen3-VL-8B-Thinking models. The token limit includes all generated reasoning and answer text. Image resolution is specified as a pixel budget rather than a fixed shape.
Figure 7: Structured self-explanation prompt used in the diagnostic analysis. It elicits six reasoning components, an alternative check, and a final answer while retaining the original task context and visual inputs.
Type
Primary
Primary (%)
Multi-label
Multi-label (%)
E1: Visual observation
59
36.42
88
54.32
E2: Physical mapping
86
53.09
121
74.69
E3: Evidence integration
5
3.09
61
37.65
E4: Pairwise comparison
1
0.62
8
4.94
E5: Factor disentanglement
10
6.17
42
25.93
E6: Target alignment
1
0.62
2
1.23
Appendix
Table 7: Error attribution for 162 selected Gemini-3-Flash failures, judged by GPT-5.5. Primary counts include only the highest-ranked error per case, and multi-label counts include all assigned errors.
Reasoning component
Mean score
Visual observation
0.90
Physical mapping
0.68
Evidence integration
0.35
Pairwise comparison
0.67
Factor disentanglement
1.39
Target alignment
1.91
Appendix
Table 8: Scores for 162 selected Gemini-3-Flash failures. Entries are means over individual cases. Higher scores indicate better satisfaction of each criterion. Each component is rated 0–2, and appropriately non-applicable comparison and disentanglement components receive 2 points.
Task
Norm.
Task
Norm.
C01
85.00
S05
28.00
C02
70.00
S06
100.00
C03
100.00
S07
100.00
C04
45.71
S08
92.00
C05
73.33
O01
100.00
C06
26.00
O02
80.00
Appendix
Table 11: Per-task chance-normalized ViT accuracy on PhysFieldBench. Total is the mean across all 24 tasks.
Fraction
Training examples
Covered
Held-out
All
10%
450
2.18
0.17
1.68
20%
900
2.96
7.33
4.06
50%
2,250
52.59
25.83
45.90
100%
4,500
58.15
16.17
47.65
Appendix
Table 12: Answer-only SFT data scaling. Fractions are relative to the maximum training size of 4,500 examples. Scores are chance-normalized accuracy averaged over the eighteen training-covered tasks, six held-out tasks, or all 24 tasks. Each row represents one training run.
Task
Zero-shot
Two-shot
Δ
C01
42.50
47.50
+5.00
C03
70.00
50.00
-20.00
C04
17.14
11.43
-5.71
C05
56.67
46.67
-10.00
C06
24.00
28.00
+4.00
C08
16.67
13.33
-3.33
Appendix
Table 13: Zero-shot and two-shot Gemini 3 Flash chance-normalized scores on training-covered tasks. Δ is Two-shot minus Zero-shot in normalized-score points. Mean is the average across tasks.
Model
C01
C02 †
C03
C04
C05
C06
C07 †
C08
Mean
Supervised visual baseline
ViT
85.0
70.0
100.0
45.7
73.3
26.0
25.0
13.3
54.8
Zero-shot MLLMs
Qwen3-VL-8B-Instruct
0.0
0.0
-20.0
-2.9
10.0
14.0
-15.0
3.3
-1.3
LLaVA-OneVision-2-8B
7.5
0.0
-5.0
-2.9
3.3
-2.0
0.0
0.0
0.1
Intern-S1-Pro
25.0
28.0
10.0
-5.7
3.3
4.0
50.0
0.0
14.3
Appendix
Table 14: Chance-normalized scores on controlled equation fields. Mean averages C01–C08. † marks tasks excluded from MLLM post-training. ViT uses per-task supervision. Within the zero-shot MLLMs, column-best scores are shown in bold . Within the Qwen3-VL-8B-Thinking variants, column-best scores are underlined . Ties are all marked.
Model
S01
S02
S03
S04
S05
S06
S07 †
S08 †
Mean
Supervised visual baseline
ViT
80.0
42.2
100.0
68.0
28.0
100.0
100.0
92.0
76.3
Zero-shot MLLMs
Qwen3-VL-8B-Instruct
0.0
0.0
-30.0
-20.0
-8.0
4.0
4.0
16.0
-4.3
LLaVA-OneVision-2-8B
2.2
0.0
0.0
-20.0
-12.0
-9.3
0.0
-20.0
-7.4
Intern-S1-Pro
17.8
15.6
-30.0
4.0
20.0
-4.0
48.0
32.0
12.9
Appendix
Table 15: Chance-normalized scores on simulated physical fields. Mean averages S01–S08. † marks tasks excluded from MLLM post-training. ViT uses per-task supervision. Within the zero-shot MLLMs, column-best scores are shown in bold . Within the Qwen3-VL-8B-Thinking variants, column-best scores are underlined . Ties are all marked.
Model
O01
O02
O03 †
O04
O05
O06
O07 †
O08
Mean
Supervised visual baseline
ViT
100.0
80.0
95.0
97.8
64.0
84.0
72.0
48.0
80.1
Zero-shot MLLMs
Qwen3-VL-8B-Instruct
0.0
-10.0
35.0
0.0
4.0
28.0
52.0
0.0
13.6
LLaVA-OneVision-2-8B
0.0
0.0
0.0
0.0
8.0
4.0
0.0
0.0
1.5
Intern-S1-Pro
-5.0
-5.0
15.0
6.7
8.0
24.0
24.0
44.0
14.0
Appendix
Table 16: Chance-normalized scores on observed physical fields. Mean averages O01–O08. † marks tasks excluded from MLLM post-training. ViT uses per-task supervision. Within the zero-shot MLLMs, column-best scores are shown in bold . Within the Qwen3-VL-8B-Thinking variants, column-best scores are underlined . Ties are all marked.
Figure 8: Representative examples from the three physical-field domains. (a) C01: spacetime fields illustrating four dominant mechanisms, with time increasing downward. (b) S06: paired NASA-CRM pressure and skin-friction fields for joint Mach-number and angle-of-attack comparison. (c) O08: paired satellite observations for radar-derived strong-VIL coverage comparison. Mechanism labels in (a) identify the examples for illustration and are not supplied to the evaluated model.
Figure 9: Successful physical-field reasoning by Gemini 3 Flash. In cooling-time comparison (S03), density growth relative to the initial state supports the correct cooling-time ordering. In cyclone-size comparison (O05), cloud-band and precipitation extent jointly support the correct size ordering.
Figure 10: Failed physical-field reasoning by Gemini. In nonlinear-advection comparison (C04), misread propagation directions lead to reversing the coefficient signs and magnitude ordering. In shear-flow comparison (S02), the model mistakes periodic structures for chaotic breakdown and associates tracer blurring directly with lower Sc.
Figure 11: Exact evaluation prompt for C01. Given a single 1D spacetime field, the model identifies the dominant equation term from five candidate mechanisms. The coefficient-based definition in the prompt specifies the candidate-selection criterion; all retained cases additionally undergo the mechanism-specific manual verification described in Table 5 .
Figure 12: Exact evaluation prompt for S06. Given two NASA-CRM cases visualized through multi-view pressure-coefficient and skin-friction fields, the model jointly compares their Mach numbers and angles of attack.
Figure 13: Exact evaluation prompt for O08. Given paired VIS, IR069, and IR107 satellite observations, the model identifies the scene with larger current strong-VIL coverage. The radar-derived VIL field is used only to construct the label and is not provided as input.