Harness Learning Enables Generalizable Test-Time Adaptation
Abstract
A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at hand. We introduce harness learning, which trains a proposer model to revise a solver's harness using execution feedback. We formulate this process as meta-learning over executable programs, with harness revisions playing the role of weight updates in gradient-based adaptation. We train the proposer with reinforcement learning, using the task performance of revised harnesses as the reward. At test time, the proposer uses feedback from successive executions on a new task to refine the harness, without performing any parameter-space update. Experiments on reasoning and multi-hop question answering show that harness learning improves revision quality and that the ability to adapt at test time transfers to unseen tasks. Policies trained on individual revisions can continue improving harnesses over multiple rounds, while the benefits of training on revision sequences vary across settings. These findings suggest a path towards continually learning agents that turn accumulated experience into generalizable improvements.
Figures & tables
| Reasoning Gym | Multi-hop QA | |
|---|---|---|
| Models | Qwen3.5-4B proposer and solver | Qwen3-4B proposer, Qwen3-8B solver |
| Training | SFT on 35B-teacher revisions, then RL; Multistep RL continues from Single-step RL | RL from Base; Multistep RL trains on online revision runs |
| Tasks | 21 SFT and 5 RL families; 21 unseen families | Train on HotpotQA; MuSiQue and 2WikiMultihopQA unseen |
| Revision | 8 independent proposals on format-varied questions; 5 rounds of one proposal, kept only if development score improves | 80 independent proposals; 4 runs of 10 rounds, 8 proposals per round |
| Reported | Held-out score; best of eight as an oracle | Held-out exact match; best of as an oracle |
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Work | Training | Adaptation |
|---|---|---|
| Harness-R1 | Teacher SFT followed by GRPO | Edit a frozen target agent’s executable harness from failure feedback. |
| JIT-Agent | Teacher SFT, execution-based DPO, supervised repair, and Evo-GDPO | Generate, repair, and evolve task-conditioned harnesses using an expanding archive. |
| Ornith | Joint RL for scaffolding and solving; Ornith-1.5 also trains task generation | Improve harness generation together with the task-solving policy. |
| Harness Learning | SFT then RL on Reasoning Gym; direct RL on QA | Revise a current harness with a fixed solver; study transfer and repeated adaptation. |
| Setting | Value |
|---|---|
| Development instances | 25 |
| Held-out instances | 75 |
| Independent proposals, single-step | 8 |
| Revision rounds, multistep | 5 |
| Proposals per revision round | 1 |
| Maximum generation attempts per proposal | 3 |
| Phase | Families |
|---|---|
| SFT | advanced_geometry , binary_alternation , bitwise_arithmetic , boxnet , calendar_arithmetic , color_cube_rotation , count_bits , countdown , dice , fraction_simplification , jugs , knights_knaves , largest_island , maze , mini_sudoku , prime_factorization , products , simple_equations , syllogism , tower_of_hanoi , tsumego |
| RL | codeio , maze , rotten_oranges , simple_geometry , tower_of_hanoi |
| Setting | Value |
|---|---|
| Data collection | |
| Teacher | Qwen3.6-35B-A3B |
| Development instances per context | 25 (16 generated for largest_island and tsumego ) |
| Failed examples in the report | Up to 8, with question text and trace excerpts |
| Teacher proposals per context | 8 |
| Retained examples | 432 |
| Family | Seed parent | Stronger parent | Total |
|---|---|---|---|
| tower_of_hanoi | 100 | 100 | 200 |
| codeio | 100 | 100 | 200 |
| rotten_oranges | 100 | 0 | 100 |
| simple_geometry | 100 | 0 | 100 |
| maze | 40 | 0 | 40 |
| Total | 440 | 200 | 640 |
| Setting | Value |
|---|---|
| Proposer sampling and optimization | |
| Initialization | SFT checkpoint |
| Candidates per context ( ) | 8 |
| Contexts per update | 40 |
| Learning rate | |
| Updates per sampled batch | 1 |
| Stage | Procedure |
|---|---|
| Initialization | Start from the reported Single-step RL checkpoint (update 10). |
| First phase | Train for 9 updates on 62 stored parent harnesses with room for improvement. |
| State refresh | Advance each parent to its best sampled revision when it improves by at least 0.01, rescore all states on the same 25 development questions, and keep states at least 0.10 below their family’s best score. |
| Second phase | Train for 15 updates on 246 states that pass the same filter, most drawn from earlier revision runs. The final checkpoint is reported as Multistep RL. |
| Family | Exposure | Tool loop | Base | SFT | Single-step RL | Single-step RL (later) | Teacher |
|---|---|---|---|---|---|---|---|
| color_cube_rotation | SFT | 0.75 | 0.19 / 0.80 | 0.75 / 0.84 | 0.60 / 0.84 | 0.76 / 0.82 | 0.32 / 1.00 |
| countdown | SFT | 0.55 | 0.30 / 0.62 | 0.42 / 0.68 | 0.38 / 0.70 | 0.55 / 0.66 | 0.28 / 1.00 |
| codeio | RL | 0.32 | 0.24 / 0.29 | 0.17 / 0.29 | 0.27 / 0.34 | 0.30 / 0.36 | – |
| rotten_oranges | RL | 0.45 | 0.39 / 1.00 | 0.21 / 0.44 | 0.45 / 1.00 | 0.44 / 1.00 | 1.00 / 1.00 |
| simple_geometry | RL | 0.53 | 0.40 / 0.68 | 0.34 / 0.56 | 0.48 / 0.58 | 0.55 / 0.57 | 0.83 / 1.00 |
| maze | SFT+RL | 0.95 | 0.05 / 0.36 | 0.68 / 0.93 | 0.86 / 0.95 | 0.87 / 0.97 | 0.97 / 0.99 |
| Evaluation set | Method | Attempted | Failed edits | Zero-scoring harnesses | Combined rate |
|---|---|---|---|---|---|
| Canonical | Base | 120 | 1 | 45 | 38% |
| SFT | 120 | 0 | 27 | 23% | |
| Single-step RL | 120 | 0 | 18 | 15% | |
| Single-step RL (later) | 120 | 0 | 14 | 12% | |
| Format-varied | Base | 96 | 4 | 29 | 34% |
| SFT | 96 | 0 | 21 | 22% |
| Format-varied (11 families) | Unseen (19 families) | |||||||
|---|---|---|---|---|---|---|---|---|
| Method | No edit | Seed | Zero | Applied | No edit | Seed | Zero | Applied |
| Base | 0 | 0.254 | 0.254 | 0.254 | 5 | 0.304 | 0.274 | 0.304 |
| SFT | 0 | 0.415 | 0.415 | 0.415 | 0 | 0.400 | 0.400 | 0.400 |
| Single-step RL | 0 | 0.576 | 0.576 | 0.576 | 0 | 0.501 | 0.501 | 0.501 |
| Single-step RL (later) | 0 | 0.551 | 0.551 | 0.551 | 1 | 0.526 | 0.519 | 0.525 |
| Multistep RL | 1 | 0.534 | 0.528 | 0.535 | ||||
| Selection group | Criteria |
|---|---|
| First six families | Single-step RL outperforms SFT, which outperforms Base; the best of eight RL proposals reaches the fixed tool-loop baseline; the seed leaves room for improvement. |
| Remaining six | The performance-ordering requirement is relaxed to both trained proposers outperforming Base. |
| Canonical | Format-varied | |||
|---|---|---|---|---|
| Method | Score | Gain over Base | Score | Gain over Base |
| Base | 0.593 | 0.498 | ||
| SFT | 0.698 | 0.688 | ||
| Single-step RL | 0.699 | 0.725 | ||
| Multistep RL | 0.701 | 0.649 | ||
| Family | Exposure | Seed | Base | SFT | Single-step RL | Multistep RL | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| r1 | r5 | r1 | r5 | r1 | r5 | r1 | r5 | |||
| basic_arithmetic | OOD | 0.49 | 0.56 | 0.58 | 0.49 | 0.92 | 0.49 | 0.87 | 0.79 | 0.91 |
| bf | OOD | 0.00 | 0.00 | 0.00 | 0.04 | 0.04 | 0.00 | 0.00 | 0.00 | 0.04 |
| color_cube_rotation | SFT | 0.17 | 0.29 | 0.33 | 0.80 | 0.88 | 0.80 | 0.80 | 0.71 | 0.82 |
| count_primes | OOD | 0.16 | 0.16 | 1.00 | 1.00 | 1.00 | 0.96 | 1.00 | 0.97 | 0.97 |
| countdown | SFT | 0.57 | 0.57 | 0.57 | 0.57 | 0.72 | 0.57 | 0.60 | 0.67 | 0.68 |
| Family | Exposure | Seed | Base | SFT | Single-step RL | Multistep RL | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| r1 | r5 | r1 | r5 | r1 | r5 | r1 | r5 | |||
| basic_arithmetic | OOD | 0.43 | 0.43 | 0.43 | 0.80 | 0.93 | 0.84 | 0.92 | 0.58 | 0.86 |
| codeio | RL | 0.28 | 0.28 | 0.28 | 0.33 | 0.35 | 0.29 | 0.40 | 0.31 | 0.38 |
| color_cube_rotation | SFT | 0.25 | 0.25 | 0.41 | 0.45 | 0.84 | 0.70 | 0.74 | 0.70 | 0.74 |
| count_primes | OOD | 0.08 | 0.68 | 0.72 | 0.32 | 0.94 | 0.87 | 0.98 | 0.96 | 0.96 |
| countdown | SFT | 0.54 | 0.54 | 0.73 | 0.61 | 0.65 | 0.63 | 0.70 | 0.54 | 0.69 |
| Setting | Composition-prompt RL | Harder-task RL |
|---|---|---|
| Training pool | 600 contexts combining low-scoring cases from the single-step RL pool ( Appendix B.3 ) with additional data; includes 100 maze and 100 sokoban contexts | 694 contexts from seven families listed below |
| Prompt | Helper-composition directive | Canonical questions without a composition directive |
| Reward | Eq. 6 | Eq. 9 , with revised validity terms and a structural-composition bonus |
| Training updates | 16 | 14 |
| Evaluated checkpoints | 10 and 16 | 14 |
| Single-step (75 held-out) | Round 5 (25 development) | |||
|---|---|---|---|---|
| Method | Format-varied | 19-family set | Canonical | Format-varied |
| Base | 0.304 | 0.496 | 0.398 | |
| SFT | 0.463 | 0.400 | 0.645 | 0.625 |
| Single-step RL | 0.506 | 0.501 | 0.660 | 0.649 |
| Single-step RL (later) | 0.526 | |||
| Multistep RL | 0.534 | 0.637 | 0.618 | |
| Method | Structural | Functional | Most common class | |
| Unseen families, without a directive | ||||
| SFT | 39 | 0 | 0 | Interpreter loop (30) |
| Single-step RL | 40 | 0 | 0 | Interpreter loop (31) |
| Unseen families, composition directive | ||||
| SFT | 24 | 11 | 4 | Tool loop + helper (12) |
| Single-step RL | 32 | 12 | 3 | Interpreter loop (17) |
| Method | Without directive | With directive |
|---|---|---|
| Single-step RL | 0.506 | 0.461 |
| Composition-prompt RL (update 10) | 0.526 | 0.559 |
| Method | Checkpoint | Round-five score |
|---|---|---|
| SFT | Initialization | 0.645 |
| Single-step RL | 10 | 0.660 |
| Tool-use reward | 14 | 0.624 |
| Combined reward | 12 | 0.624 |
| Combined reward | 20 | 0.630 |
| Component | Observed failure | Consequence |
|---|---|---|
| Question parsing | The parser misses the goal marker. | Questions are routed to a fallback model call. |
| Answer extraction | The harness returns an unsolved grid. | The submitted answer is incomplete. |
| Tool communication | A tool message is malformed. | The tool interaction cannot execute. |
| Benchmark | Retrieval corpus | Dev pool | Test | Split source |
|---|---|---|---|---|
| HotpotQA | 5.23M Wikipedia abstracts (full-wiki) | 500 | 300 | train split, disjoint subsets |
| MuSiQue | 21,100 paragraphs pooled from the dev questions | 500 | 299 | dev split, hop-stratified |
| 2WikiMultihopQA | 54,957 paragraphs pooled from the dev questions | 500 | 300 | dev split, type-stratified |
| Setting | Value |
|---|---|
| Contexts | |
| Parent harness | Seed (Single-step); state of a revision run advanced by the policy (Multistep) |
| Feedback / scoring questions | 6 / 64, disjoint, from a 2,000-question training pool |
| Revision runs depth (Multistep) | 32 10; commit by best-of-8 scoring exact match |
| Contexts per phase | 320 |
| Proposer sampling and optimization | |
| Dataset | Training | Seed | Mean | Oracle@8 | Best-of-80 | seed | Failed |
|---|---|---|---|---|---|---|---|
| HotpotQA | Single-step RL | 0.3900 | 0.4105 | 0.5399 | 0.5800 | 71.3% | 1.3% |
| Multistep RL | 0.3900 | 0.4131 | 0.5209 | 0.5833 | 77.5% | 1.3% | |
| MuSiQue | Single-step RL | 0.1438 | 0.1467 | 0.2126 | 0.2408 | 58.8% | 5.0% |
| Multistep RL | 0.1438 | 0.1551 | 0.2034 | 0.2408 | 67.5% | 2.5% | |
| 2WikiMultihopQA | Single-step RL | 0.2500 | 0.3090 | 0.4746 | 0.5200 | 67.5% | 1.3% |
| Multistep RL | 0.2500 | 0.2748 | 0.4374 | 0.5433 | 58.8% | 0.0% |