MedicalHarness: A Controlled Evaluation of LLMs and Agent Harnesses on Medical Tasks
Organizations: Northwestern University
Abstract
LLM agents are increasingly built for medical work and scored on clinical benchmarks. Each such score, however, comes from a model running inside an agent harness, the system that controls the loop between the model and its environment. An agent's score is therefore a property of a model--harness pair. For medical agents, how much outcomes change with the harness has rarely been measured. Measuring this change, and explaining it, raises two challenges. First, a harness comparison must change nothing but the harness and be repeated across models and kinds of task. Second, comparing whole harnesses leaves their mechanisms bundled together, so it cannot show when an individual mechanism helps. To address these challenges, we present MedicalHarness, a controlled study of models and agent harnesses on medical tasks. We first build MedicalHarnessBench to evaluate agents on tasks across four domains that each test a different harness capability. Using this benchmark, we run five open-weight models under five agent harnesses, changing only the harness within a comparison, and analyze both outcomes and execution traces. To study individual mechanisms, we build MH-Lab, a controlled harness that switches off context management, planning or tool exposure one at a time within a shared execution loop. We find that the harness and its interaction with the model account for about a quarter of the outcome variance, and that no single harness is best across models and tasks. Code and data are available at https://github.com/REAL-Lab-NU/MedicalHarness.
Figures & tables
| Harness | Rank | ||||||||||||
| Outc. | Proc. | MTok | Outc. | Proc. † | MTok | Outc. | Proc. | MTok | Outc. | Proc. | MTok | ||
| ZeroClaw | 0.08 0.01 | 0.37 0.01 | 0.11 | 0.49 0.03 | 0.32 0.01 | 0.48 | 0.01 0.02 | 0.20 0.02 | 1.79 | 0.03 0.02 | 0.20 0.05 | 6.67 | 3.75 |
| OpenClaw | 0.10 0.01 | 0.45 0.04 | 0.52 | 0.57 0.08 | 0.35 0.05 | 1.04 | 0.00 0.00 | 0.20 0.01 | 4.51 | 0.04 0.04 | 0.38 0.06 | 11.95 | 3.00 |
| Hermes | 0.13 0.04 | 0.46 0.05 | 0.11 | 0.57 0.12 | 0.23 0.07 | 0.78 | 0.00 0.00 | 0.32 0.04 | 1.04 | 0.01 0.02 | 0.15 0.13 | 4.02 | 3.75 |
| Codex | 0.09 0.01 | 0.47 0.04 | 0.06 | 0.77 0.05 | 0.48 0.07 | 0.73 | 0.01 0.02 | 0.31 0.04 | 2.30 | 0.11 0.02 | 0.29 0.05 | 20.04 | 2.25 |
| Condition | ||||||||
| Outc. | MTok | Outc. | MTok | Outc. | MTok | Outc. | MTok | |
| full | 0.10 0.02 | 0.23 0.01 | 0.88 0.03 | 0.43 0.12 | 0.12 0.05 | 1.62 0.26 | 0.44 0.02 | 5.96 0.38 |
| context management | 0.01 0.02 | 1.1 0.0 | 0.12 0.03 | 0.5 0.2 | 0.04 0.06 | 0.4 0.1 | 0.39 0.09 | 0.3 0.0 |
| planning | 0.02 0.04 | 1.0 0.1 | 0.04 0.04 | 0.8 0.4 | 0.01 0.08 | 0.9 0.2 | 0.00 0.07 | 0.7 0.2 |
| tool exposure | 0.00 0.02 | 3.0 1.6 | 0.14 0.05 | 2.4 0.7 | 0.01 0.05 | 2.6 0.5 | 0.00 0.07 | 0.8 0.1 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Env. | Source (license) | Pool | Selection (tasks left) | Kept | Verification |
| MedMCP-Calc (MIT) | difficulty screen ( ) reliability contract and independent model audit ( ) | scorer accepts the gold and rejects a perturbed answer | |||
| MedMemoryBench (Apache-2.0) | exact-match answers ( ) long history ( ) answer not restated in recent sessions ( ) entity and date questions whose gold the record supports ( ) two blind model reviews ( , plus from an earlier round) | reviewers lock answers before the gold is compared | |||
| HealthAgentBench (MIT) | runs without credentialed data ( ) | upstream task tests | |||
| HealthAdminBench (Apache-2.0) | – outcome checks and a process check, fixed quota per family ( ) | upstream deterministic checks |
| Environment | Display group | Detailed content | |
| ( tasks) | Internal Medicine | Gastroenterology/hepatology ( ), infectious diseases ( ), respiratory medicine ( ), hematology/oncology ( ), cardiovascular medicine ( ), and nephrology/electrolytes ( ). | |
| Acute & Perioperative Care | Emergency medicine/trauma ( ), intensive care ( ), and anesthesia/perioperative care ( ). | ||
| Neurology & Mental Health | Neurosurgery/neurology ( ) and mental/psychological health ( ). | ||
| Nursing & Pediatrics | Inpatient nursing-risk assessment ( ) and pediatric-asthma assessment ( ). | ||
| ( tasks) | Test Results | Retrieve recorded laboratory, imaging, or measured examination findings, including changes in measured indices. | |
| Medication History | Retrieve drug identities or classes, doses, dose changes, discontinuation, or the identity of an allergenic medication. |
| Env. | Example task (paraphrased) | Best |
| A patient with ulcerative colitis is admitted with frequent bloody stools and systemic symptoms. Decide which severity and risk calculators the question calls for, pull their inputs from the patient’s record, compute them, and write every value to the answer file. | ||
| Given consultation sessions in order, answer one question about a dose, lab value or diagnosis that was recorded once, dozens of sessions before the question. | ||
| Given one admission note and a directory of trial documents, list every trial whose inclusion criteria the patient meets and none of whose exclusion criteria apply. Other tasks select the tumor region on a whole-slide image. | ||
| For a no-authorization claim denial, review the denial, the remittance image and the patient inquiry in the EMR, check eligibility on the payer portal, carry out the right resolution there, and document it in a triage note. |
| Env. | Outcome | Answer read from | Diagnostic (Proc.) |
| calculator-selection F1 share of required outputs correct, with omitted outputs counted wrong. Numbers match within of the reference , and categories after normalization | answer file, else final reply | calculator-selection F1 | |
| if the normalized answer matches an accepted alias | answer file, else final reply | LLM judge on the condensed trace | |
| if the task’s own tests pass | output files | reference-set recall ( tasks) | |
| if every outcome check on the final state passes | application state | share of process checks passed |
| Harness | Version | Default changed, and why |
| OpenClaw | 2026.9.2 | output cap K (truncated tool calls on of tasks), clock s s |
| Hermes | 0.21.0 | none |
| ZeroClaw | 0.8.5 | tool-iteration limit (past it, the tool list is dropped) |
| Codex CLI | 0.156.1 | API-key variable supplied, MCP tools auto-approved ( ) |
| Claude Code | 2.1.267–2.1.270 | none (the proxy translates its Anthropic-format traffic) |
| Mean | |||||
| Model | |||||
| Harness | |||||
| Interaction | |||||
| Run-to-run | |||||
| HV/MV, five models | |||||
| HV/MV, three main models |
| Harness | Rank | ||||||||||||
| Outc. | Proc. | MTok | Outc. | Proc. † | MTok | Outc. | Proc. | MTok | Outc. | Proc. | MTok | ||
| ZeroClaw | 0.05 0.02 | 0.23 0.04 | 0.53 | 0.55 0.09 | 0.20 0.07 | 0.88 | 0.01 0.02 | 0.17 0.10 | 3.23 | 0.00 0.00 | 0.08 0.02 | 9.26 | 4.50 |
| OpenClaw | 0.06 0.01 | 0.28 0.07 | 0.65 | 0.59 0.05 | 0.24 0.05 | 1.35 | 0.00 0.00 | 0.15 0.01 | 5.00 | 0.00 0.00 | 0.08 0.02 | 12.22 | 3.75 |
| Hermes | 0.07 0.00 | 0.36 0.02 | 0.23 | 0.62 0.05 | 0.31 0.05 | 0.85 | 0.02 0.04 | 0.40 0.08 | 1.87 | 0.00 0.00 | 0.07 0.03 | 11.54 | 2.50 |
| Codex | 0.08 0.01 | 0.33 0.03 | 0.60 | 0.68 0.05 | 0.41 0.05 | 0.72 | 0.02 0.02 | 0.25 0.09 | 2.95 | 0.04 0.04 | 0.08 0.01 | 8.01 | 1.50 |
| Model | ||||||
| Leader | Gap [95% CI] | Leader | Gap [95% CI] | |||
| Hermes | 0.95 | 0.04 [0.00, 0.09] | Claude Code | 0.61 | 0.01 [ 0.07, 0.10] | |
| Hermes | 0.91 | 0.02 [ 0.01, 0.04] | Codex | 0.47 | 0.00 [ 0.10, 0.10] | |
| Claude Code | 0.94 | 0.02 [0.00, 0.04] | Claude Code | 0.97 | 0.12 [ 0.03, 0.23] | |
| Codex | 0.80 | 0.01 [ 0.01, 0.03] | Claude Code | 0.68 | 0.03 [ 0.06, 0.12] | |
| Hermes | 0.59 | 0.00 [ 0.01, 0.02] | Codex | 0.66 | 0.06 [ 0.09, 0.20] | |
| Model | Answer rate | Score (shared) | Answer rate | Score (shared) | ||
| 0.50–0.95 | 0.12–0.16 | 28 | 0.57–0.91 | 0.87–0.90 | 30 | |
| 0.46–1.00 | 0.10–0.16 | 39 | 0.81–1.00 | 0.82–0.88 | 56 | |
| 0.80–0.98 | 0.13–0.17 | 57 | 0.59–1.00 | 0.56–0.85 | 27 | |
| 0.59–0.81 | 0.08–0.12 | 28 | 0.77–0.99 | 0.69–0.73 | 45 | |
| 0.91–0.98 | 0.02–0.04 | 72 | 0.80–0.96 | 0.44–0.75 | 32 | |
| Setting | Value |
| Matrix | two models, four conditions, tasks and three seeds, for scheduled episodes |
| Input budget | tokens per request for both models, with reserved for token-estimation error |
| Window and output cap | : K window, K output cap. : K window, K output cap |
| File tools | seven workspace tools for reading, writing, editing, listing and searching files and running shell commands, with workspace confinement and read-before-write checks |
| Browser tools | fourteen browser tools |
| Execution limits | tool results capped at K characters, loop detection, a -step limit and a one-hour deadline |
| Condition | ||
| Outc. [95% CI] | Outc. [95% CI] | |
| full | 0.10 0.02 | 0.88 0.03 |
| context management | 0.01 [ 0.01, 0.04] | 0.12 [ 0.23, 0.01] |
| planning | 0.02 [ 0.04, 0.00] | 0.04 [ 0.12, 0.01] |
| tool exposure | 0.00 [ 0.02, 0.03] | 0.14 [ 0.23, 0.07] |
| context management | planning | tool exposure | ||||
| Environment | Overflows | Median turns | No valid submission | Timeouts | Rejected calls per episode | Episodes with a rejected call |
| 0 0 | 10 9 | 24 47 | 22 44 | 18 | 95/96 | |
| 0 8 | 9 6 | 1 0 | 1 0 | 43 | 66/69 | |
| 2 57 | 37 38 | 3/84 3/81 | 8/84 16/81 | 79 | 84/84 | |
| 0 67 | 99.5 78.5 | – | 3 1 | 31 | 47/72 | |
| Module | Harness model | Environment, window | Outcome | Ovf. | Pairs |
| Summarization | OpenClaw | , 80K | 69 | ||
| Summarization | Deep Agents | , 80K | 69 | ||
| Summarization | Deep Agents | , 64K | 69 | ||
| Summarization | Hermes | , 80K | 23 | ||
| Summarization | Deep Agents | , 80K | 69 | ||
| Tool-output truncation | Hermes | , native | – | 23 |