OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
Authors: Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li, Baiding Chen, Zijun Dou, Qingcheng Zeng, Qi Kang, +23 more
Organizations: University of California, Los Angeles · RIKEN AIP · Yale University · University of Waterloo · Carnegie Mellon University · Tsinghua University · Northwestern University · Zhejiang University · University of California, Berkeley · University of Illinois Chicago · Boston University · The University of Hong Kong · Stanford University · The University of Tokyo · New York University
Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorld-Science, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain. The benchmark contains 12 VLMs and 146 high-quality tasks across several scientific domains and software configurations, covering workflows such as molecular drawing and retrosynthesis, pathology image analysis, statistical computing, and physical simulation. Tasks are developed through expert proposals and iterative human--AI co-design, with selection guided by scientific value and difficulty. Task-specific execution-based evaluators inspect application states and generated artifacts, including molecular structures, segmentation masks, plots, and numerical results, and award partial credit for incomplete outcomes. Our special harness integrates model adapters, interaction-loop control, and trajectory logging to support comparisons of models and interaction strategies. Our results show that current state-of-the-art VLMs with a strong harness still face challenges in addressing key questions in the scientific domains. We also analyze the benchmarking results across multi-linguistics, reasoning efforts, context length and other factors and derive several important conclusions and directions to assist future development. Overall, we provide an integrated framework connecting expert-defined scientific goals to verifiable software outcomes, enabling systematic evaluation of both agent capabilities and harness design in scientific workflows.
Figures & tables
Figure 1: Landscape of task collection, environment setup, harness construction, and evaluation pipelines of OSWorld-Science.
Figure 2: Composition of the collected OSWorld-Science task set. (a) Task shares by scientific domain. (b) Task counts by recorded software configuration. Panels (a) and (d) are normalized by the full task set. Slash and plus labels preserve paired configurations in the source inventory.
Figure 3: OSWorld-Science example workflow and overall performance-cost comparison. (a) An agent uses ASKCOS to address a scientific task. (b) Mean task-specific score, including partial credit and with VOID runs scored zero, versus mean API cost per task in U.S. dollars (logarithmic scale) across 12 evaluated VLMs.
Figure 4: Trajectory analysis of the 1,530 released runs. (a) Outcome by model. (b) Per-run means (%) of the interaction measures and of steps containing four common automation calls ( † chemistry and runs without recorded thinking excluded); shading is normalized per column. (c) Full-credit minus other runs: median over tasks (circles; 86 tasks, 69 for thinking) or over model-domain groups (diamonds; 39 groups, 34 for thinking), with bootstrap 95% intervals.
Figure 5: Ablations on medical subset, the 23 QuPath tasks; one run per cell, VOID runs scored 0, output tokens include thinking tokens. Top left: reasoning-effort sweep at w=5 . Top right: history-window sweep for Opus 5 at medium effort. Bottom: task-prompt language at medium effort and w=5 ; the English bars are the medium cells of the effort sweep. No cell differs significantly from its default under a paired sign-flip test after Bonferroni correction over the 22 non-default cells (smallest uncorrected p=0.04 , Opus 5 with French prompts); lines follow sweep order and are visual guides only.
Figure 6: Contrasting Opus 5 chemistry workflows with representative raw trajectory frames shown directly below each workflow diagram. (a) and (b): the successful Lenacapavir run moves from source-table localization to focused stereochemical inspection and final validation of the complete 17-row CSV. (c) and (d): the failed Archangiumide run converts coordinate-guided reconstruction into a chemically valid, internally consistent structure, yet its own round-trip check still reveals a stereochemical mismatch. The decisive difference is independent source-grounded validation rather than trajectory length or the number of checks.
Appendix figures & tables45 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Overall performance versus output-token consumption across the 12 evaluated VLMs. The vertical axis shows mean task score with VOID runs scored zero; the horizontal axis shows mean output tokens per task among tasks with recorded usage. Token consumption and API dollar cost are distinct efficiency measures.
Figure 8: Overall mean API cost per task in U.S. dollars, ordered from highest to lowest. The bars provide the cost values underlying the overall performance–cost comparison.
Figure 9: Overall mean task score across the 12 evaluated VLMs, ordered from highest to lowest. Scores include task-specific partial credit, with VOID runs assigned zero; they are not binary task-completion rates.
Figure 10: Performance versus API cost overall and by domain. Each point represents a model; the vertical axis shows mean task score with VOID runs scored zero, and the horizontal axis shows mean API cost per task in U.S. dollars on a logarithmic scale. Panel titles report task counts. Only model–domain pairs with available score and cost values are plotted. Domain panels use different axis ranges, and labels round scores to whole percentages.
Opus 5
Sonnet 5
Prompt language
Score
n
English
Prompt lang.
Other
Score
n
English
Prompt lang.
Other
English
44.1
65
100.0
100.0
0.0
15.7
1629
100.0
100.0
0.0
Chinese
29.7
55
10.9
89.1
0.0
10.9
1604
99.9
0.1
0.0
Japanese
53.1
51
35.3
64.7
0.0
18.0
1660
82.2
17.8
0.0
Spanish
43.0
28
57.1
42.9
0.0
11.7
1498
95.2
4.8
0.0
French
33.9
32
90.6
9.4
0.0
11.3
1516
98.2
1.8
0.0
Appendix
Table 1: Multilingual experiments of the model replies in the task-language sweep (medium effort, w=5 ). Score: mean partial-credit score (%), as in Table 2 . n : number of replies over the 23 runs that contain natural-language text outside code blocks and control tokens; the remaining columns are the percentages of these n replies written in English, in the language of the prompt, or in any other language (for English prompts the first two coincide). Opus 5 emits such prose in only 28-65 replies per cell, so its shares rest on few replies.
Opus 5
Sonnet 5
Setting
Partial
Binary
Steps
#Tok.
Partial
Binary
Steps
#Tok.
Reasoning effort ( w=5 , English prompts)
low
42.7
34.8
22
6.3
15.0
8.7
69
18.8
medium
44.1
34.8
27
13.8
15.7
13.0
87
23.8
high
50.3
43.5
43
29.7
22.0
17.4
83
35.4
xhigh
47.7
34.8
42
44.8
13.5
8.7
85
36.5
Appendix
Table 2: Ablation results on QuPath-Bench (23 QuPath tasks; one run per cell, step limit 100). Partial: mean partial-credit score (%). Binary: percentage of tasks that receive a score of 1 (one task corresponds to 4.3 points). Steps: mean interaction steps per task. #Tok.: mean output tokens per task in thousands, thinking tokens included. VOID runs count as 0 under both score metrics. Shaded rows are the shared default configuration (medium effort, w=5 , English prompts), a single run that appears in every block. Within each block, the best value per column is in bold and the second best is underlined (lower is better for Steps and #Tok.; ties share the mark).
Setting
Partial
Binary
Steps
#Tok.
Tok./step
Think
GUI
CLI
Prose
Opus 5
Reasoning effort ( w=5 , English prompts)
low
42.7
34.8
22
6.3
281
0.40
0.20
0.84
0.08
medium
44.1
34.8
27
13.8
510
0.57
0.37
0.84
0.15
high
50.3
43.5
43
29.7
699
0.68
0.45
0.78
0.16
xhigh
47.7
34.8
42
44.8
1061
0.70
0.46
0.83
0.67
Appendix
Table 3: Performance and trajectory statistics for every ablation cell (23 QuPath tasks, one run per cell, step limit 100). Partial, Binary, Steps, and #Tok. as in Table 2 . Tok./step: mean output tokens per task divided by mean steps. Think: share of output tokens that are thinking tokens. GUI and CLI: share of action steps that contain a mouse action or a terminal/script command, respectively (a step can count as both). Prose: share of model replies that contain at least 20 characters of natural-language text outside code blocks and control tokens. Shaded rows are the shared default configuration (medium effort, w=5 , English prompts), a single run per backbone that appears in every block.
Opus 5
Sonnet 5
Setting
Score
Solved
Partial
Zero
VOID
Score
Solved
Partial
Zero
VOID
Reasoning effort ( w=5 , English prompts)
low
42.7
8
7
6
2
15.0
2
4
11
6
medium
44.1
8
8
5
2
15.7
3
4
8
8
high
50.3
10
4
5
4
22.0
4
7
6
6
xhigh
47.7
8
6
6
3
13.5
2
6
11
4
Appendix
Table 4: Outcome composition of every ablation cell. Score: mean partial-credit score (%), as in Table 2 ; the remaining columns count tasks out of 23. Solved: score 1. Partial: score strictly between 0 and 1. Zero: score 0 with a gradable final state. VOID : no gradable outcome. Shaded rows are the shared default configuration.
Opus 5
Sonnet 5
Task
low
medium
high
xhigh
max
low
medium
high
xhigh
max
QP_001_T1-1
1.00
0.00
0.00
0.00
1.00
0.00
0.00
0.00
0.00
0.00
QP_001_T1-2
1.00
1.00
1.00
0.00
1.00
0.00
0.00
1.00
0.00
1.00
QP_001_T2-1
0.08
0.03
0.00
0.00
0.10
0.00
V
V
0.00
V
QP_001_T2-2
0.27
0.23
V
V
V
0.00
V
V
V
V
QP_001_T3
0.00
0.00
0.58
0.30
0.63
0.45
V
0.17
0.00
0.20
Appendix
Table 5: Per-task scores in the reasoning-effort sweep ( w=5 , English prompts). V: VOID run (no gradable outcome). The bottom rows give the mean score with VOID scored 0, the number of fully solved tasks, and the number of VOID runs.
Task
w =1
w =3
w =5
w =10
w =15
QP_001_T1-1
0.00
0.00
0.00
0.00
0.00
QP_001_T1-2
1.00
1.00
1.00
1.00
1.00
QP_001_T2-1
0.15
0.09
0.03
0.00
0.00
QP_001_T2-2
0.17
V
0.23
0.24
0.31
QP_001_T3
0.07
0.30
0.00
0.38
0.00
QP_002_T1
0.00
1.00
0.00
1.00
1.00
Appendix
Table 6: Per-task scores in the history-window sweep (Opus 5, medium effort, English prompts). V: VOID run (no gradable outcome). The bottom rows give the mean score with VOID scored 0, the number of fully solved tasks, and the number of VOID runs.
Opus 5
Sonnet 5
Task
en
zh
ja
es
fr
th
en
zh
ja
es
fr
th
QP_001_T1-1
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
QP_001_T1-2
1.00
0.00
0.00
0.00
1.00
0.00
0.00
0.00
1.00
0.00
0.00
0.00
QP_001_T2-1
0.03
0.12
V
0.00
0.00
0.11
V
V
V
V
V
V
QP_001_T2-2
0.23
0.00
0.17
0.00
0.22
0.00
V
V
V
V
V
V
QP_001_T3
0.00
0.34
0.32
0.00
0.16
0.60
V
V
V
V
V
0.18
Appendix
Table 7: Per-task scores in the task-language sweep (medium effort, w=5 ). V: VOID run (no gradable outcome). The bottom rows give the mean score with VOID scored 0, the number of fully solved tasks, and the number of VOID runs.
Figure 11: Outcome of each run by software configuration, all models pooled (12 models, 11 on QGIS; number of runs in parentheses); categories as in Figure 4 a.
No deliverable
Software
Runs
Full
Wrong
Invalid
Placeh.
None
Harness
Unres.
ND/fail (%)
QuPath
276
48
89
3
0
134
2
0
59
Radiol.
36
8
10
1
1
16
0
0
61
ANSYS
168
66
5
5
0
91
1
0
90
OpenFOAM
168
77
17
8
2
64
0
0
73
CIAO
36
14
13
0
0
8
1
0
38
Appendix
Table 8: Run outcomes by software configuration (12 models; counts of runs; 11 models on QGIS). Placeh.: the only artifact is a placeholder without answer content; None: no artifact, or the task-provided file left unchanged. Harness: ended by the harness or environment without an artifact; Unres.: the released records do not determine the artifact state. ND/fail: runs without a deliverable as a percentage of the remaining runs that did not receive full credit. Radiol.: Weasis and 3D Slicer; CIAO: CIAO+DS9; Chem.: ChemDraw, ASKCOS, and PDF-viewer tasks.
Model
QuPath
Radiol.
ANSYS
OpenFOAM
CIAO
SAS/R
Chem.
QGIS
Praat
All
Claude Fable 5.1
48
42
77
28
69
25
35
80
100
44
Claude Opus 5
37
90
64
37
74
64
35
94
98
49
GPT-6 Astra
65
49
71
20
40
37
15
77
98
40
GPT-5.6 sol
76
70
87
50
73
12
24
–
90
45
Kimi K3
40
82
76
8
68
25
45
82
88
44
GPT-5.6 terra
83
68
77
22
69
4
35
89
93
47
Appendix
Table 9: GUI steps by model and software configuration: percentage of action steps that contain a pointer action (click, move, drag, scroll), averaged over the runs with at least one action. – : no runs (GPT-5.6 sol has no QGIS runs). Radiol.: Weasis and 3D Slicer; CIAO: CIAO+DS9; Chem.: ChemDraw, ASKCOS, and PDF-viewer tasks.
Model
QuPath
Radiol.
ANSYS
OpenFOAM
CIAO
SAS/R
Chem.
QGIS
Praat
All
Claude Fable 5.1
69
56
37
74
54
93
47
24
0
59
Claude Opus 5
84
71
40
61
72
74
47
49
0
59
GPT-6 Astra
38
40
22
78
76
61
23
51
0
40
GPT-5.6 sol
23
30
17
84
19
60
55
–
0
47
Kimi K3
45
10
11
78
24
49
40
37
0
41
GPT-5.6 terra
22
26
17
78
21
75
62
13
0
48
Appendix
Table 10: CLI steps by model and software configuration: percentage of action steps that open a terminal or contain a shell command or a call to the application’s scripting interface, averaged over runs. The values are lower bounds (see text); CLI steps are below 1% in Praat, whose tasks are solved in the GUI. Radiol.: Weasis and 3D Slicer; CIAO: CIAO+DS9; Chem.: ChemDraw, ASKCOS, and PDF-viewer tasks.
GUI operations (% of steps)
Model
Click
Dbl.
Right
Drag
Move
Scroll
Type
Press
Hotkey
Narr.
Think.
Rep.
Steps
Claude Fable 5.1
34
7
0
0
2
5
45
8
12
65
46
3
12
Claude Opus 5
44
2
0
0
1
1
65
9
11
29
60
1
18
GPT-6 Astra
27
11
0
2
6
3
43
16
16
2
17
3
14
GPT-5.6 sol
37
7
0
1
4
2
48
53
23
1
40
4
21
Kimi K3
36
3
1
1
10
4
54
14
10
69
56
7
48
Appendix
Table 11: Interaction profile by model, averaged over runs in all domains. Operation columns: percentage of action steps that contain at least one call of that type (a step may contain several). Narr.: percentage of replies with narration; Think.: thinking tokens as a percentage of output tokens (chemistry and runs without recorded thinking excluded, see text); Rep.: percentage of replies whose action repeats one of the previous three. Steps: median steps of runs with full credit.
Figure 12: Delivered geometries (yellow) on two digitization tasks. Columns from left to right: the input image, Opus 5, Fable 5.1, and GPT-6-Astra. Top row: building footprints (GEO_004). The scene contains 4 large buildings and 8 small sheds beside the parking lots, and only GPT-6-Astra digitizes most of the sheds. Bottom row: road centerlines (GEO_005). Opus 5 and GPT-6-Astra cover the full network, and Fable 5.1 leaves part of it out.
Cause
Criterion
Runs
Backbones (runs)
Action does not take effect
A
Coordinate-frame mismatch
Click coordinates do not follow the 1920×1080 pixel frame stated in the instruction, so clicks miss their target
18
Qwen3.7-plus (6), Gemini 3.1 Pro (5), Qwen-CUA (4), MiniMax-M3 (3)
B
No executable action
Reply contains no executable action, for example only a text plan
4
Kimi-K3 (3), Qwen-CUA (1)
C
Malformed action
Actions are malformed and never execute
1
MiniMax-M3 (1)
Artifact is malformed
D
Data type or geometry error
Wrong field type or geometry dimension, so the result is invalid
3
Gemini 3.1 Pro, MiniMax-M3, Qwen-CUA (1 each)
Appendix
Table 12: Primary causes of the 53 failed runs on the geoscience tasks, ordered along the execution path. Each failed run has exactly one cause. The last column lists the affected backbones with their number of runs.
Backbone
GEO_001
GEO_002
GEO_003
GEO_004
GEO_005
GEO_006
Pass
Score
GPT-6-Astra
P
P
E
F
P
E
3/6
69.0
Fable 5.1
P
P
E
F
F
E
2/6
68.1
Opus 5
P
P
E
F
P
H
3/6
61.6
GPT-5.6-luna
P
G
E
F
E
E
1/6
48.6
Sonnet 5
P
G
E
F
E
H
1/6
37.1
GPT-5.6-terra
G
P
E
F
F
E
1/6
36.5
Appendix
Table 13: Outcome of every run on the six geoscience tasks (one run per cell, step limit 100). P: successful run. A-H: primary cause of failure as in Table 12 , shaded by stage. Pass: number of successful runs. Score: mean partial-credit score in percent.
model
runs
shell
app.
applications driven
declined
mean
med.
med. input
GUI
aloud
score
steps
tokens (k)
gpt-6-astra
20
13
7
SAS Studio
0
0.898
12
158
kimi-k3
20
14
6
SAS Studio, VS Code, gedit
3
0.698
38
2310
gemini-3.1-pro
20
16
4
RStudio, VS Code, gedit
0
0.730
11
105
minimax-m3
20
16
4
VS Code, gedit
8
0.667
68
1348
claude-sonnet-5
20
17
3
VS Code
4
0.774
40
745
Appendix
Table 14: Applications used by each model across its 20 statistics tasks. shell denotes a run that opened none of SAS Studio, RStudio, gedit, or VS Code. Such a run may have used a terminal window, executed code without a window, or done both. app. GUI denotes a run that drove at least one of the four applications. declined aloud denotes a run that named an application and argued for using the shell in the same reply. Median input tokens are computed per run and include screenshots resent at every step.
Figure 13: Software routes across all statistics runs. Five of the twelve models never opened SAS Studio, RStudio, gedit, or VS Code in their twenty tasks. File-viewer use is concentrated in the five tasks with raster inputs.
task group
runs
shell
shell +
app.
obstacle
mean
only
viewer
GUI
only
score
SAS offered (9 tasks)
108
93
2
13
0
0.938
RStudio named (2 tasks)
24
18
0
6
0
0.571
raster inputs (5 tasks)
60
29
27
4
2
0.726
R only (4 tasks)
48
43
2
3
0
0.871
all
240
183
31
26
2
0.835
Appendix
Table 15: The same census by task group. shell + viewer denotes a shell run that also opened an image or PDF viewer, as required by the five tasks with raster inputs. obstacle only denotes a run in which the desktop or a stray keystroke placed an application in front of the agent, which then spent steps closing it. These cases do not count as application use.
Figure 14: The 108 runs on the nine tasks that offer SAS OnDemand for Academics. The guest starts with the sign-in page displayed and the credentials saved, so agents did not need to find or launch the browser. Bars are drawn to scale. The final stage contains one run, kimi-k3 on task stat_liver_cohort .
outcome (means)
comparison
n shell
n app.
shell
app.
difference
score
raw
214
26
0.861
0.625
-0.236
same task
within task
142
26
-0.139 ( p=0.017 )
same model
within model
114
26
-0.299 ( p<0.001 )
steps used
raw
214
26
28
58
30
same task
within task
142
26
29 ( p<0.001 )
same model
within model
114
26
22 ( p=0.002 )
Appendix
Table 16: Outcomes and costs by interaction mode. Raw rows compare all shell-mode and application-mode runs. Each stratified row computes the application-minus-shell difference within tasks or models that contain both modes, then averages the differences weighted by stratum size. The number of shell-mode runs is therefore smaller in these rows because strata without an application-mode run do not contribute. The p -values come from 20,000 permutations of interaction-mode labels within strata. All entries are means.
Figure 15: Outcomes and costs by application type. Each point represents one run, and the horizontal marker shows the median. Steps and tokens use logarithmic axes. Using SAS Studio or RStudio is associated with more steps but no detectable score difference. Using a GUI text editor is associated with both more steps and lower scores. The two runs that used both a domain application and an editor are assigned to the domain-application group.
Figure 16: Two interaction modes for stat_group_tests . Each panel shows the run’s own screenshot, with the relevant region boxed and reproduced below at native resolution. (a) gpt-6-astra has signed in and selected Launch, but SAS Studio is still loading. The run instead inspects the local files in a terminal and commits to R at step 7, earning a score of 1.000. (b) kimi-k3 has uploaded both data files, run its macro, and read the correct Wilcoxon output from the rendered page. Its budget ends three steps later with an empty submission directory. The score is 0.080, awarded only for the protected-input gate.
neg
plosive1
Backbone
bun
bat
Pat
bit
pit
Score
Steps
5 ms
2 ms
5 ms
2 ms
5 ms
neg / plos.
neg / plos.
gpt-6-astra
+3.1
−0.1
−0.5
−0.1
−2.2
1.00 / 1.00
9 / 27
claude-fable-5-1
+0.6
+4.1
+4.9
+4.4
+3.2
1.00 / 0.55
15 / 47
gpt-5.6-sol
done, no file
+1.9
+0.6
+2.1
+2.8
0.10 / 0.78
71 / 23
kimi-k3
−1.6
no file
1.00 / 0.10
75 / 100
Appendix
Table 17: Outcome of every boundary on the two linguistics tasks (one run per cell, step limit 100). Numeric entries report signed placement error in milliseconds, computed as the placed boundary minus the reference. Scores include the 0.1 awarded for returning the audio unchanged.
Figure 17: The four voicing onsets in plosive1 , showing the reference and the boundaries from the three runs that placed all four. The shaded region marks the tolerance, and T is the local glottal period. Within each run, all four signed errors have the same direction. On bat and bit , voicing begins with a small pulse followed by a larger one. claude-fable-5-1 marks the start of the larger pulse.
Figure 18: Three runs locating the release of bun in neg . Each row pairs the run’s Praat view with an enlargement of the boxed waveform region. Where visible, the black line marks the reference release, the red line marks the inserted boundary, and the grey triangle marks the earlier transient within the prevoicing. (a) gpt-6-astra rejects the transient and places the boundary 3.1 ms after the reference, within tolerance. Its score is 1.000. (b) claude-opus-5 inspects a 16 ms window through the Select dialog but marks the transient 5.5 ms before the reference. Its score is 0.10. (c) The window viewed by claude-sonnet-5 begins 6 ms after the reference. The run treats the edge of its pink selection as the release and places the boundary 51.7 ms late. Its score is 0.10.
Figure 19: EEG analysis and event-pairing audit by Opus 5 on two tasks. (a) Epoch and analysis windows for trial-aligned RT–P3 inference. (b) Mean Pz amplitudes over 300–500 ms for 36 fast and 36 slow trials; horizontal bars indicate group means. (c) Actual event timing for the first four stimuli in the forensic task. The candidate crosses a trial boundary and reuses a response preceding a later stimulus; the corrected mapping leaves unmatched stimuli unanswered. (d) Six shared response events affect 12 stimulus rows, whereas the agent reports six surplus assignments. Right: original terminal observations with relevant output highlighted.
Cause
Criterion
Runs
Backbones (runs)
Action does not take effect
A Action misses
Clicks follow a coordinate frame other than the 1920 × 1080 screen, or input goes to another window, so the step has no effect
Mnova is never installed, so no spectrum is opened
1
Gemini 3.1 Pro
Method or content is wrong
C Query wrong
The database query excludes valid entries or never runs
2
Qwen3.7-plus, MiniMax-M3
Appendix
Table 18: Primary causes of the 27 failed runs on the structural-biology and NMR tasks, ordered along the execution path. Each failed run has exactly one cause. The last column lists the affected backbones with their number of runs.
Backbone
align
af-egfr
af-her2
af-igf1r
mutation
rcsb
synergy
nmr
Pass
Score
Sonnet 5
P
P
P
P
P
P
P
E (0.75)
7/8
96.9
Opus 5
P
P
P
P
P
P
P
E (0.75)
7/8
96.9
Fable 5.1
P
P
P
P
P
P
P
E (0.72)
7/8
96.5
GPT-6-Astra
P
P
P
P
P
P
F V
E (0.69)
6/8
83.6
GPT-5.6-luna
P
P
P
P
P
P
A V
A V
6/8
75.0
GPT-5.6-sol
P
P
P
P
P
P
A V
A V
6/8
75.0
Appendix
Table 19: Outcome of every run on the eight structural-biology and NMR tasks (one run per cell, step limit 100). P: successful run. A–F: primary cause of failure as in Table 18 , shaded by stage with the colors of Table 13 . A superscript V marks a VOID run, which reaches the step limit without DONE, submits nothing, and scores 0. For the partially credited nmr runs the score is given in parentheses. Pass: number of successful runs. Score: mean partial-credit score in percent over the scored tasks.
Figure 20: The final screen of the seven nmr runs that never open a spectrum (one run per backbone). Six click Mnova’s license agreement and Registration Wizard in a scaled coordinate frame (A), and Gemini 3.1 Pro never installs Mnova (B). They end on the Registration Wizard, the license file chooser, the license agreement, the vendor’s web store, the vendor’s 45-day trial form, a shell reporting command not found , or the launch command typed into a help window, with the institutional license files in /home/user/licenses/ throughout. Institution names are redacted.
Figure 21: Composition of the collected OSWorld-Science task set-continued. (a) Shares of 266 non-exclusive required-capability annotations; screenshot-based visual observation is required for all 145 tasks. (b) Availability of CLI support proportion.
Figure 22: Contrasting workflows on QP_004_T1 (register an attention map onto a reference slide; requires installing the Interactive Image Alignment extension). Top: GPT-6 Astra installs through the application, confirms the menu entry exists, applies the affine transform, and switches to the command finder when menus stall — 31 steps, score 1.0. Bottom: Claude Opus 5 copies the jar into a guessed directory from the shell and relaunches QuPath nine times. Panels (c2) and (c3) show the same empty Analyze menu at step 19 and step 98: the disconfirming evidence was available 79 steps before the budget ran out. The run is scored VOID . The decisive difference is verifying the precondition through the interface that owns it, not trajectory length or effort.
Figure 23: Contrasting workflows on nmr_6bromoindole_assignment (assign the 1 H spectrum of 6-bromoindole from the raw FID in Mnova; requires accepting Mnova’s license agreement, choosing Install in its Registration Wizard and importing a license file). Top: Claude Opus 5 installs Mnova, clicks Accept and Install where they are, and Mnova itself confirms the license (a2) before the run analyzes the spectrum (a3) — 25 steps, score 0.75. Bottom: GPT-5.6-luna clicks one point beside the Registration Wizard 39 times; scaled by 1.41, the point is the Install button (red cross: the click; green circle: Install). Panels (c1) and (c2) show the same unchanged wizard at step 50 and step 96; the run turned to the keyboard only at step 97, three steps before the budget ran out. The run is scored VOID . The decisive difference is whether the clicks land in the screen’s coordinate frame, not knowledge of NMR. Institution names in the license file names are redacted.
Figure 24: Contrasting GPT-6-Astra and GPT-5.6-terra workflows on netcounts , with representative raw trajectory frames shown beside each workflow diagram. (a) and (b): the successful run reads its regions back from DS9. It then reads NET_COUNTS from the output file named at the bottom of the Net Counts window (green boxes) and writes 1424.9204178903. (c) and (d): the failed run sees the same window with the NET_COUNTS row scrolled out of view. It enters 1427−BG_ERR=1410.7212 (red boxes) and later cites its own entry as the correct value. The decisive difference is the origin of the reported number, not the setup, the route, or the trajectory length.
Cause
Criterion
Runs
Backbones (runs)
Action does not take effect
A
Coordinate-frame mismatch
Click coordinates do not follow the 1920×1080 frame stated in the system prompt, so clicks land away from the intended terminal, menu item, or dialog button
Region radii given without the arcsecond mark are read as pixels
1
MiniMax-M3 (1)
Result is misread
C
Visible number taken as the result
A number shown on the screen, or the difference of two such numbers, is entered as the background-subtracted value
2
GPT-5.6-luna, GPT-5.6-terra (1 each)
Appendix
Table 20: Primary causes of the 22 failed runs on the astronomy tasks, ordered along the execution path. Each failed run has exactly one cause. The last column lists the affected backbones with their number of runs.
Backbone
netcounts
lightcurve
streak
Pass
Score
Opus 5
P
P
P
3/3
100.0
Fable 5.1
P
P
P
3/3
100.0
GPT-5.6-sol
P
P
E
2/3
66.7
GPT-6-Astra
P
P
F
2/3
66.7
Kimi-K3
G V
P
P
2/3
66.7
Sonnet 5
P
D (0.5)
E
1/3
50.0
Appendix
Table 21: Outcome of every run on the three astronomy tasks (one run per cell). P: successful run. A–G: primary cause of failure as in Table 20 , shaded by stage. A superscript V marks a VOID run, which reaches the step limit without DONE , submits nothing, and scores 0. For each failed lightcurve run that delivers an answer, the partial-credit score is given in parentheses. Pass: number of successful runs. Score: mean partial-credit score in percent.
Cause
Criterion
Runs
Backbones (runs)
Action does not take effect
A 1
Coordinate-frame mismatch
Clicks follow a scaled frame (normalized or 1280×720 ) instead of the 1920×1080 frame stated in the system prompt, so launcher and dialog clicks miss their target
39
Gemini 3.1 Pro (14), Qwen3.7-plus (14), MiniMax-M3 (6), Qwen-CUA (5)
A 2
Input lost in a widget
Typed text lands in the console’s find box or is dropped by the Export dialog’s quantity filter, so commands and selections never take effect
Replies contain no executable action until the stall limit
1
MiniMax-M3 (1)
C
Malformed action
Replies chain hundreds of identical action blocks, so no step makes progress
1
Qwen-CUA (1)
Artifact is malformed
Appendix
Table 22: Primary causes of the 102 ANSYS runs without full credit (99 failed and 3 partially credited runs), ordered along the execution path. Each run has exactly one cause. The last column lists the affected backbones with their number of runs.
FT
FV
Backbone
005c
006b
008a
008c
001a
001b
002a
002b
003a
003b
004a
004b
005a
005b
Pass
Score
Fable 5.1
P
P
P
P
P
P
P
P
P
P
P
P
P
P
14/14
100.0
GPT-6-Astra
P
P
P
P
H
P
P
P
P
P
H
P
P
P
12/14
85.7
Opus 5
P
P
P
P
P
P
H
P
P
P
P
P
H
P
12/14
85.7
GPT-5.6-sol
D
H
P
P
P
P
P
P
P
P
P
H
H
P
10/14
71.4
Kimi-K3
P
H
P
P
H
P
H
H
P
H
H
P
P
P
8/14
57.1
Appendix
Table 23: Outcome of every run on the 14 ANSYS Fluent tasks (one run per cell, step limit 100; nine successful runs used an earlier limit of 60 to 80 steps). P: successful run. A–H: primary cause of failure as in Table 22 , shaded by stage; a superscript gives the partial-credit score in percent. Pass: number of successful runs. Score: mean partial-credit score in percent.
Figure 25: Contrasting two trajectories on the adjoint task FT-008a, with raw frames from each run shown to the right of its workflow diagram. (a) and (b): Kimi-K3 evaluates its drag observable, reads 0 N, selects the wall zone it had left unselected, re-evaluates to 1271.7444 N, and runs the adjoint to convergence at iteration 27 (score 1.0, 86 steps). (c) and (d): GPT-5.6-terra reaches the same dialog with the same 0 N reading, concludes that the observable is correctly configured, obtains 1271.7444 N from the stored drag report without reconciling the two values, and runs a zero-source adjoint that stops at iteration 1 (score 0.33, 100 steps). Each frame is the run’s own screenshot with the relevant region boxed. The decisive difference is whether an implausible value may overrule the agent’s reading of its setup, not the trajectory length or the number of checks.
Figure 26: Contrasting GPT-5.6-sol OpenFOAM workflows, with representative raw terminal frames beside each workflow diagram (cropped to the terminal window; a step number refers to the screen the model saw before its action at that step). (a) and (b): the successful flat-plate run (OF-012) inspects the array declarations in its VTK ( .vtu ) output, rejects the Float32 arrays as short of full precision (step 33), reads the 12-digit native field files (step 34), and checks row counts and physical bounds before EXPORT_OK (score 1 after 35 steps). (c) and (d): the failed plane-channel run (OF-004) sees the same kind of declaration (step 37), builds its CSV files from 6-digit ASCII VTK arrays and checks their line counts, array lengths, and header (step 38), and declares DONE although its own volume-weighted bulk velocity differs from Ub in the 8th significant figure (step 41; score 0.125 after 41 steps). The string FFPP in (c) lists the gates G1–G4 in order, F for fail and P for pass. Green boxes mark the evidence that the successful run acted on, and red boxes mark what the failed run accepted and delivered. The decisive difference is validation of the delivered precision against the stated requirement, not the correctness of the solve or the number of checks.
Cause
Criterion
Runs
Backbones (runs)
Action does not take effect
A
Input never reaches the case
Keystrokes land outside the terminal, the action code fails before typing (a syntax error from nested quoting), or the agent’s own here-document closes around nothing, and the lost case or solver is never re-sent
The agent issues FAIL or DONE with steps left, before a solver has run on every required mesh
4
Gemini 3.1 Pro (2), GPT-5.6-luna, MiniMax-M3 (1 each)
C
Solver never launched
The step limit ends the run before a solver is launched on at least one required mesh, and this missing launch, not a start-up error, is the decisive defect; the steps go to reading source code, tutorials or the agent’s own files
Every launch on at least one required mesh stops at a start-up error before the first iteration (missing or headerless dictionary, missing scheme entry, wrong patch type), and the run ends with such an error unresolved
Table 24: Primary causes of the 91 failed runs on the OpenFOAM tasks, ordered along the execution path. Each failed run has exactly one cause: the defect that decides its outcome, and the earliest one along the path when there are several. The path is judged per mesh: a run with a mesh that never iterates gets one of the causes A–D, with two exceptions. The cause is E when the delivered matrices come from a substitute for the simulation (four runs), and F when the agent never reached that mesh because a wrong setting derailed the solve on the other mesh and was never fixed (five runs). The last column lists the affected backbones with their number of runs.
Backbone
001
002
004
005
007
009
010
011
012
013
015
016
017
018
Pass
Score
Fable 5.1
P
P
P
P
P
P
P
P
P
P
P
P
P
P
14/14
100.0
GPT-6-Astra
P
P
P
P
P
P
P
P
P
P
P
F
P
P
13/14
93.8
GPT-5.6-sol
P
P
I
P
P
P
I
P
P
P
P
P
P
F
11/14
82.1
Opus 5
P
P
P
P
P
P
P
P
G
P
G V
F
G
P
10/14
72.3
Kimi-K3
P
C V
H V
D V
P
P
P
D V
P
P
F
C V
C V
P
7/14
50.9
Gemini 3.1 Pro
P
D V
B
F
P
P
P
H
B
P
G
F
P
H V
6/14
43.8
Appendix
Table 25: Outcome of every run on the 14 OpenFOAM tasks (one run per cell, step limit 100). Column headers give the task number, for example 001 for OF-001. P: successful run. A–I: primary cause of failure as in Table 24 , shaded by stage. A superscript V marks a VOID run, which reaches the step or wall-clock limit without DONE , submits nothing, and scores 0. Pass: number of successful runs. Score: mean partial-credit score in percent.
Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems, whereas scientific instrumentation scenarios require coordinated control over complex interfaces, and feedback-driven parameter adjustment. However, directly evaluating agents on physical high-precision instruments is impractical due to high cost, safety risks, limited accessibility, and difficulty in ensuring reproducible evaluation. This motivates the need for a simulated yet realistic testbed that preserves the operational challenges of scientific instruments while enabling scalable and safe benchmarking. To this end, we introduce LabOSBench, a challenging benchmark for multimodal GUI agents built on a suite of web-based scientific-instrument simulators. Operating directly via a browser, LabOSBench avoids resource-heavy OS virtualization while supporting flexible task configuration and execution-based evaluation. Specifically, LabOSBench constructs 96 subtasks across eight instrument simulators, covering workflows from sample loading, alignment, parameter tuning, and data acquisition to result inspection. We evaluate general-purpose vision-language models, specialized GUI agent models, and advanced agentic frameworks at both subtask and end-to-end levels. Our experiments reveal that while existing agents can complete many structured GUI subtasks, they still struggle with feedback-driven operations and long-horizon workflow execution. Overall, LabOSBench provides a reproducible, low-cost testbed for advancing computer-using agents toward scientific-instrument control.
Anqi Zou, Han Deng, Chengyu Zhang +9
Shenzhen Loop Area Institute · Dalian University of Technology · The Chinese University of Hong Kong +1
Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents. We introduce OSWorld 2.0, a benchmark of 108 long-horizon computer-use workflows across everyday and professional tasks, designed to capture complex and challenging real-world phenomena. Each task represents a realistic end-to-end workflow that takes human users a median of about 1.6 hours to complete and requires an average of 318 tool calls with Claude Opus 4.7 using maximum thinking, compared with about 30 in OSWorld 1.0. OSWorld 2.0 targets challenge phenomena that are common in real workflows yet underrepresented in prior benchmarks, spanning interaction-design challenges such as streaming interaction and dynamic environments, as well as agent-pattern challenges such as cross-source reasoning, implicit-state inference, and visual-spatial precision. Tasks are grounded in authentic input artifacts and cross-referenced against realistic stateful user profile data, and include separate safety reports auditing safety-sensitive execution. Under our primary binary-completion metric at 500 steps, Claude Opus 4.8 with maximum thinking and batched tool calls scores best but still completes only 20.6% of tasks at a 54.8% partial score; GPT-5.5 is far more token-efficient yet plateaus near 13%. These results show that current agents are still far from professional-level computer use: rather than stumbling on basic GUI control or coding, they lose track of constraints, miss information that arrives mid-task, guess rather than ask the user, and skip verification, struggling most when a task hinges on hidden state they must recover.
Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visualization (SciVis) tasks. Despite rapid progress, the community lacks a principled and reproducible benchmark for evaluating these emerging SciVis agents in realistic, multi-step analysis settings. We present SciVisAgentBench, a comprehensive and extensible benchmark for evaluating scientific data analysis and visualization agents. Our benchmark is grounded in a structured taxonomy spanning four dimensions: application domain, data type, complexity level, and visualization operation. It currently comprises 108 expert-crafted cases covering diverse SciVis scenarios. To enable reliable assessment, we introduce a multimodal outcome-centric evaluation pipeline that combines LLM-based judging with deterministic evaluators, including image-based metrics, code checkers, rule-based verifiers, and case-specific evaluators. We also conduct a validity study with 12 SciVis experts to examine the agreement between human and LLM judges. Using this framework, we evaluate representative SciVis agents and general-purpose coding agents to establish initial baselines and reveal capability gaps. SciVisAgentBench is designed as a living benchmark to support systematic comparison, diagnose failure modes, and drive progress in agentic SciVis. The benchmark is available at https://scivisagentbench.github.io/.
Kuangshi Ai, Haichao Miao, Kaiyuan Tang +13
University of Notre Dame · Lawrence Livermore National Laboratory · University of Utah +4