OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
Organizations: University of California, Los Angeles · RIKEN AIP · Yale University · University of Waterloo · Carnegie Mellon University · Tsinghua University · Northwestern University · Zhejiang University · University of California, Berkeley · University of Illinois Chicago · Boston University · The University of Hong Kong · Stanford University · The University of Tokyo · New York University
Abstract
Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorld-Science, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain. The benchmark contains 12 VLMs and 146 high-quality tasks across several scientific domains and software configurations, covering workflows such as molecular drawing and retrosynthesis, pathology image analysis, statistical computing, and physical simulation. Tasks are developed through expert proposals and iterative human--AI co-design, with selection guided by scientific value and difficulty. Task-specific execution-based evaluators inspect application states and generated artifacts, including molecular structures, segmentation masks, plots, and numerical results, and award partial credit for incomplete outcomes. Our special harness integrates model adapters, interaction-loop control, and trajectory logging to support comparisons of models and interaction strategies. Our results show that current state-of-the-art VLMs with a strong harness still face challenges in addressing key questions in the scientific domains. We also analyze the benchmarking results across multi-linguistics, reasoning efforts, context length and other factors and derive several important conclusions and directions to assist future development. Overall, we provide an integrated framework connecting expert-defined scientific goals to verifiable software outcomes, enabling systematic evaluation of both agent capabilities and harness design in scientific workflows.
Figures & tables
Appendix figures & tables45 assets
Supplementary material from the paper’s appendix.
Appendix
| Opus 5 | Sonnet 5 | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Prompt language | Score | English | Prompt lang. | Other | Score | English | Prompt lang. | Other | ||
| English | 44.1 | 65 | 100.0 | 100.0 | 0.0 | 15.7 | 1629 | 100.0 | 100.0 | 0.0 |
| Chinese | 29.7 | 55 | 10.9 | 89.1 | 0.0 | 10.9 | 1604 | 99.9 | 0.1 | 0.0 |
| Japanese | 53.1 | 51 | 35.3 | 64.7 | 0.0 | 18.0 | 1660 | 82.2 | 17.8 | 0.0 |
| Spanish | 43.0 | 28 | 57.1 | 42.9 | 0.0 | 11.7 | 1498 | 95.2 | 4.8 | 0.0 |
| French | 33.9 | 32 | 90.6 | 9.4 | 0.0 | 11.3 | 1516 | 98.2 | 1.8 | 0.0 |
| Opus 5 | Sonnet 5 | |||||||
| Setting | Partial | Binary | Steps | #Tok. | Partial | Binary | Steps | #Tok. |
| Reasoning effort ( , English prompts) | ||||||||
| low | 42.7 | 34.8 | 22 | 6.3 | 15.0 | 8.7 | 69 | 18.8 |
| medium | 44.1 | 34.8 | 27 | 13.8 | 15.7 | 13.0 | 87 | 23.8 |
| high | 50.3 | 43.5 | 43 | 29.7 | 22.0 | 17.4 | 83 | 35.4 |
| xhigh | 47.7 | 34.8 | 42 | 44.8 | 13.5 | 8.7 | 85 | 36.5 |
| Setting | Partial | Binary | Steps | #Tok. | Tok./step | Think | GUI | CLI | Prose |
| Opus 5 | |||||||||
| Reasoning effort ( , English prompts) | |||||||||
| low | 42.7 | 34.8 | 22 | 6.3 | 281 | 0.40 | 0.20 | 0.84 | 0.08 |
| medium | 44.1 | 34.8 | 27 | 13.8 | 510 | 0.57 | 0.37 | 0.84 | 0.15 |
| high | 50.3 | 43.5 | 43 | 29.7 | 699 | 0.68 | 0.45 | 0.78 | 0.16 |
| xhigh | 47.7 | 34.8 | 42 | 44.8 | 1061 | 0.70 | 0.46 | 0.83 | 0.67 |
| Opus 5 | Sonnet 5 | |||||||||
| Setting | Score | Solved | Partial | Zero | VOID | Score | Solved | Partial | Zero | VOID |
| Reasoning effort ( , English prompts) | ||||||||||
| low | 42.7 | 8 | 7 | 6 | 2 | 15.0 | 2 | 4 | 11 | 6 |
| medium | 44.1 | 8 | 8 | 5 | 2 | 15.7 | 3 | 4 | 8 | 8 |
| high | 50.3 | 10 | 4 | 5 | 4 | 22.0 | 4 | 7 | 6 | 6 |
| xhigh | 47.7 | 8 | 6 | 6 | 3 | 13.5 | 2 | 6 | 11 | 4 |
| Opus 5 | Sonnet 5 | |||||||||
| Task | low | medium | high | xhigh | max | low | medium | high | xhigh | max |
| QP_001_T1-1 | 1.00 | 0.00 | 0.00 | 0.00 | 1.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| QP_001_T1-2 | 1.00 | 1.00 | 1.00 | 0.00 | 1.00 | 0.00 | 0.00 | 1.00 | 0.00 | 1.00 |
| QP_001_T2-1 | 0.08 | 0.03 | 0.00 | 0.00 | 0.10 | 0.00 | V | V | 0.00 | V |
| QP_001_T2-2 | 0.27 | 0.23 | V | V | V | 0.00 | V | V | V | V |
| QP_001_T3 | 0.00 | 0.00 | 0.58 | 0.30 | 0.63 | 0.45 | V | 0.17 | 0.00 | 0.20 |
| Task | =1 | =3 | =5 | =10 | =15 |
|---|---|---|---|---|---|
| QP_001_T1-1 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| QP_001_T1-2 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| QP_001_T2-1 | 0.15 | 0.09 | 0.03 | 0.00 | 0.00 |
| QP_001_T2-2 | 0.17 | V | 0.23 | 0.24 | 0.31 |
| QP_001_T3 | 0.07 | 0.30 | 0.00 | 0.38 | 0.00 |
| QP_002_T1 | 0.00 | 1.00 | 0.00 | 1.00 | 1.00 |
| Opus 5 | Sonnet 5 | |||||||||||
| Task | en | zh | ja | es | fr | th | en | zh | ja | es | fr | th |
| QP_001_T1-1 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| QP_001_T1-2 | 1.00 | 0.00 | 0.00 | 0.00 | 1.00 | 0.00 | 0.00 | 0.00 | 1.00 | 0.00 | 0.00 | 0.00 |
| QP_001_T2-1 | 0.03 | 0.12 | V | 0.00 | 0.00 | 0.11 | V | V | V | V | V | V |
| QP_001_T2-2 | 0.23 | 0.00 | 0.17 | 0.00 | 0.22 | 0.00 | V | V | V | V | V | V |
| QP_001_T3 | 0.00 | 0.34 | 0.32 | 0.00 | 0.16 | 0.60 | V | V | V | V | V | 0.18 |
| No deliverable | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Software | Runs | Full | Wrong | Invalid | Placeh. | None | Harness | Unres. | ND/fail (%) |
| QuPath | 276 | 48 | 89 | 3 | 0 | 134 | 2 | 0 | 59 |
| Radiol. | 36 | 8 | 10 | 1 | 1 | 16 | 0 | 0 | 61 |
| ANSYS | 168 | 66 | 5 | 5 | 0 | 91 | 1 | 0 | 90 |
| OpenFOAM | 168 | 77 | 17 | 8 | 2 | 64 | 0 | 0 | 73 |
| CIAO | 36 | 14 | 13 | 0 | 0 | 8 | 1 | 0 | 38 |
| Model | QuPath | Radiol. | ANSYS | OpenFOAM | CIAO | SAS/R | Chem. | QGIS | Praat | All |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 48 | 42 | 77 | 28 | 69 | 25 | 35 | 80 | 100 | 44 |
| Claude Opus 5 | 37 | 90 | 64 | 37 | 74 | 64 | 35 | 94 | 98 | 49 |
| GPT-6 Astra | 65 | 49 | 71 | 20 | 40 | 37 | 15 | 77 | 98 | 40 |
| GPT-5.6 sol | 76 | 70 | 87 | 50 | 73 | 12 | 24 | – | 90 | 45 |
| Kimi K3 | 40 | 82 | 76 | 8 | 68 | 25 | 45 | 82 | 88 | 44 |
| GPT-5.6 terra | 83 | 68 | 77 | 22 | 69 | 4 | 35 | 89 | 93 | 47 |
| Model | QuPath | Radiol. | ANSYS | OpenFOAM | CIAO | SAS/R | Chem. | QGIS | Praat | All |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 69 | 56 | 37 | 74 | 54 | 93 | 47 | 24 | 0 | 59 |
| Claude Opus 5 | 84 | 71 | 40 | 61 | 72 | 74 | 47 | 49 | 0 | 59 |
| GPT-6 Astra | 38 | 40 | 22 | 78 | 76 | 61 | 23 | 51 | 0 | 40 |
| GPT-5.6 sol | 23 | 30 | 17 | 84 | 19 | 60 | 55 | – | 0 | 47 |
| Kimi K3 | 45 | 10 | 11 | 78 | 24 | 49 | 40 | 37 | 0 | 41 |
| GPT-5.6 terra | 22 | 26 | 17 | 78 | 21 | 75 | 62 | 13 | 0 | 48 |
| GUI operations (% of steps) | |||||||||||||
| Model | Click | Dbl. | Right | Drag | Move | Scroll | Type | Press | Hotkey | Narr. | Think. | Rep. | Steps |
| Claude Fable 5.1 | 34 | 7 | 0 | 0 | 2 | 5 | 45 | 8 | 12 | 65 | 46 | 3 | 12 |
| Claude Opus 5 | 44 | 2 | 0 | 0 | 1 | 1 | 65 | 9 | 11 | 29 | 60 | 1 | 18 |
| GPT-6 Astra | 27 | 11 | 0 | 2 | 6 | 3 | 43 | 16 | 16 | 2 | 17 | 3 | 14 |
| GPT-5.6 sol | 37 | 7 | 0 | 1 | 4 | 2 | 48 | 53 | 23 | 1 | 40 | 4 | 21 |
| Kimi K3 | 36 | 3 | 1 | 1 | 10 | 4 | 54 | 14 | 10 | 69 | 56 | 7 | 48 |
| Cause | Criterion | Runs | Backbones (runs) | |
| Action does not take effect | ||||
| A | Coordinate-frame mismatch | Click coordinates do not follow the pixel frame stated in the instruction, so clicks miss their target | 18 | Qwen3.7-plus (6), Gemini 3.1 Pro (5), Qwen-CUA (4), MiniMax-M3 (3) |
| B | No executable action | Reply contains no executable action, for example only a text plan | 4 | Kimi-K3 (3), Qwen-CUA (1) |
| C | Malformed action | Actions are malformed and never execute | 1 | MiniMax-M3 (1) |
| Artifact is malformed | ||||
| D | Data type or geometry error | Wrong field type or geometry dimension, so the result is invalid | 3 | Gemini 3.1 Pro, MiniMax-M3, Qwen-CUA (1 each) |
| Backbone | GEO_001 | GEO_002 | GEO_003 | GEO_004 | GEO_005 | GEO_006 | Pass | Score |
|---|---|---|---|---|---|---|---|---|
| GPT-6-Astra | P | P | E | F | P | E | 3/6 | 69.0 |
| Fable 5.1 | P | P | E | F | F | E | 2/6 | 68.1 |
| Opus 5 | P | P | E | F | P | H | 3/6 | 61.6 |
| GPT-5.6-luna | P | G | E | F | E | E | 1/6 | 48.6 |
| Sonnet 5 | P | G | E | F | E | H | 1/6 | 37.1 |
| GPT-5.6-terra | G | P | E | F | F | E | 1/6 | 36.5 |
| model | runs | shell | app. | applications driven | declined | mean | med. | med. input |
|---|---|---|---|---|---|---|---|---|
| GUI | aloud | score | steps | tokens (k) | ||||
| gpt-6-astra | 20 | 13 | 7 | SAS Studio | 0 | 0.898 | 12 | 158 |
| kimi-k3 | 20 | 14 | 6 | SAS Studio, VS Code, gedit | 3 | 0.698 | 38 | 2310 |
| gemini-3.1-pro | 20 | 16 | 4 | RStudio, VS Code, gedit | 0 | 0.730 | 11 | 105 |
| minimax-m3 | 20 | 16 | 4 | VS Code, gedit | 8 | 0.667 | 68 | 1348 |
| claude-sonnet-5 | 20 | 17 | 3 | VS Code | 4 | 0.774 | 40 | 745 |
| task group | runs | shell | shell + | app. | obstacle | mean |
|---|---|---|---|---|---|---|
| only | viewer | GUI | only | score | ||
| SAS offered (9 tasks) | 108 | 93 | 2 | 13 | 0 | 0.938 |
| RStudio named (2 tasks) | 24 | 18 | 0 | 6 | 0 | 0.571 |
| raster inputs (5 tasks) | 60 | 29 | 27 | 4 | 2 | 0.726 |
| R only (4 tasks) | 48 | 43 | 2 | 3 | 0 | 0.871 |
| all | 240 | 183 | 31 | 26 | 2 | 0.835 |
| outcome (means) | comparison | shell | app. | shell | app. | difference |
|---|---|---|---|---|---|---|
| score | raw | 214 | 26 | 0.861 | 0.625 | -0.236 |
| same task | within task | 142 | 26 | -0.139 ( ) | ||
| same model | within model | 114 | 26 | -0.299 ( ) | ||
| steps used | raw | 214 | 26 | 28 | 58 | 30 |
| same task | within task | 142 | 26 | 29 ( ) | ||
| same model | within model | 114 | 26 | 22 ( ) |
| neg | plosive1 | ||||||
| Backbone | bun | bat | Pat | bit | pit | Score | Steps |
| 5 ms | 2 ms | 5 ms | 2 ms | 5 ms | neg / plos. | neg / plos. | |
| gpt-6-astra | 1.00 / 1.00 | 9 / 27 | |||||
| claude-fable-5-1 | 1.00 / 0.55 | 15 / 47 | |||||
| gpt-5.6-sol | done, no file | 0.10 / 0.78 | 71 / 23 | ||||
| kimi-k3 | no file | 1.00 / 0.10 | 75 / 100 | ||||
| Cause | Criterion | Runs | Backbones (runs) |
|---|---|---|---|
| Action does not take effect | |||
| A Action misses | Clicks follow a coordinate frame other than the 1920 1080 screen, or input goes to another window, so the step has no effect | 13 | GPT-5.6-sol, GPT-5.6-luna, GPT-5.6-terra, Gemini 3.1 Pro, Qwen3.7-plus, Qwen-CUA (2 each), MiniMax-M3 |
| Blocked before the task can start | |||
| B Application not installed | Mnova is never installed, so no spectrum is opened | 1 | Gemini 3.1 Pro |
| Method or content is wrong | |||
| C Query wrong | The database query excludes valid entries or never runs | 2 | Qwen3.7-plus, MiniMax-M3 |
| Backbone | align | af-egfr | af-her2 | af-igf1r | mutation | rcsb | synergy | nmr | Pass | Score |
|---|---|---|---|---|---|---|---|---|---|---|
| Sonnet 5 | P | P | P | P | P | P | P | E (0.75) | 7/8 | 96.9 |
| Opus 5 | P | P | P | P | P | P | P | E (0.75) | 7/8 | 96.9 |
| Fable 5.1 | P | P | P | P | P | P | P | E (0.72) | 7/8 | 96.5 |
| GPT-6-Astra | P | P | P | P | P | P | F V | E (0.69) | 6/8 | 83.6 |
| GPT-5.6-luna | P | P | P | P | P | P | A V | A V | 6/8 | 75.0 |
| GPT-5.6-sol | P | P | P | P | P | P | A V | A V | 6/8 | 75.0 |
| Cause | Criterion | Runs | Backbones (runs) | |
| Action does not take effect | ||||
| A | Coordinate-frame mismatch | Click coordinates do not follow the frame stated in the system prompt, so clicks land away from the intended terminal, menu item, or dialog button | 8 | Qwen3.7-plus (3), Gemini 3.1 Pro, Qwen-CUA (2 each), MiniMax-M3 (1) |
| Measurement is set up wrong | ||||
| B | Wrong aperture unit | Region radii given without the arcsecond mark are read as pixels | 1 | MiniMax-M3 (1) |
| Result is misread | ||||
| C | Visible number taken as the result | A number shown on the screen, or the difference of two such numbers, is entered as the background-subtracted value | 2 | GPT-5.6-luna, GPT-5.6-terra (1 each) |
| Backbone | netcounts | lightcurve | streak | Pass | Score |
|---|---|---|---|---|---|
| Opus 5 | P | P | P | 3/3 | 100.0 |
| Fable 5.1 | P | P | P | 3/3 | 100.0 |
| GPT-5.6-sol | P | P | E | 2/3 | 66.7 |
| GPT-6-Astra | P | P | F | 2/3 | 66.7 |
| Kimi-K3 | G V | P | P | 2/3 | 66.7 |
| Sonnet 5 | P | D (0.5) | E | 1/3 | 50.0 |
| Cause | Criterion | Runs | Backbones (runs) | |
| Action does not take effect | ||||
| A 1 | Coordinate-frame mismatch | Clicks follow a scaled frame (normalized or ) instead of the frame stated in the system prompt, so launcher and dialog clicks miss their target | 39 | Gemini 3.1 Pro (14), Qwen3.7-plus (14), MiniMax-M3 (6), Qwen-CUA (5) |
| A 2 | Input lost in a widget | Typed text lands in the console’s find box or is dropped by the Export dialog’s quantity filter, so commands and selections never take effect | 5 | Sonnet 5 (3), GPT-5.6-luna, GPT-5.6-terra (1 each) |
| B | No executable action | Replies contain no executable action until the stall limit | 1 | MiniMax-M3 (1) |
| C | Malformed action | Replies chain hundreds of identical action blocks, so no step makes progress | 1 | Qwen-CUA (1) |
| Artifact is malformed | ||||
| FT | FV | |||||||||||||||
| Backbone | 005c | 006b | 008a | 008c | 001a | 001b | 002a | 002b | 003a | 003b | 004a | 004b | 005a | 005b | Pass | Score |
| Fable 5.1 | P | P | P | P | P | P | P | P | P | P | P | P | P | P | 14/14 | 100.0 |
| GPT-6-Astra | P | P | P | P | H | P | P | P | P | P | H | P | P | P | 12/14 | 85.7 |
| Opus 5 | P | P | P | P | P | P | H | P | P | P | P | P | H | P | 12/14 | 85.7 |
| GPT-5.6-sol | D | H | P | P | P | P | P | P | P | P | P | H | H | P | 10/14 | 71.4 |
| Kimi-K3 | P | H | P | P | H | P | H | H | P | H | H | P | P | P | 8/14 | 57.1 |
| Cause | Criterion | Runs | Backbones (runs) | |
| Action does not take effect | ||||
| A | Input never reaches the case | Keystrokes land outside the terminal, the action code fails before typing (a syntax error from nested quoting), or the agent’s own here-document closes around nothing, and the lost case or solver is never re-sent | 5 | GPT-5.6-luna (2), GPT-5.6-terra, Qwen-CUA, MiniMax-M3 (1 each) |
| No solution is computed | ||||
| B | Gives up before solving | The agent issues FAIL or DONE with steps left, before a solver has run on every required mesh | 4 | Gemini 3.1 Pro (2), GPT-5.6-luna, MiniMax-M3 (1 each) |
| C | Solver never launched | The step limit ends the run before a solver is launched on at least one required mesh, and this missing launch, not a start-up error, is the decisive defect; the steps go to reading source code, tutorials or the agent’s own files | 14 | Sonnet 5 (7), Kimi-K3 (3), MiniMax-M3 (2), Qwen3.7-plus, GPT-5.6-luna (1 each) |
| D | Launched but never iterates | Every launch on at least one required mesh stops at a start-up error before the first iteration (missing or headerless dictionary, missing scheme entry, wrong patch type), and the run ends with such an error unresolved | 26 | MiniMax-M3 (9), Qwen-CUA (6), Qwen3.7-plus (5), Kimi-K3, GPT-5.6-luna (2 each), Gemini 3.1 Pro, Sonnet 5 (1 each) |
| Backbone | 001 | 002 | 004 | 005 | 007 | 009 | 010 | 011 | 012 | 013 | 015 | 016 | 017 | 018 | Pass | Score |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Fable 5.1 | P | P | P | P | P | P | P | P | P | P | P | P | P | P | 14/14 | 100.0 |
| GPT-6-Astra | P | P | P | P | P | P | P | P | P | P | P | F | P | P | 13/14 | 93.8 |
| GPT-5.6-sol | P | P | I | P | P | P | I | P | P | P | P | P | P | F | 11/14 | 82.1 |
| Opus 5 | P | P | P | P | P | P | P | P | G | P | G V | F | G | P | 10/14 | 72.3 |
| Kimi-K3 | P | C V | H V | D V | P | P | P | D V | P | P | F | C V | C V | P | 7/14 | 50.9 |
| Gemini 3.1 Pro | P | D V | B | F | P | P | P | H | B | P | G | F | P | H V | 6/14 | 43.8 |