Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the Builder's experience reusable, we introduce Meta-Skill: principles specifying when support is needed and what resources to provide. The Builder learns these principles from Target's execution feedback on the development set, then uses the frozen skill bank to construct harnesses for unseen tasks. Across Harness-Bench and NewtonBench, full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction, and 12.02 points over direct delivery of the same bank to the Target. These results highlight the value of translating experience into executable support. Gains when the same model serves both roles further suggest a path to system level self-improvement through learning to build better environments.
Figures & tables
Figure 1 : Overview of our framework. Rather than teaching the Target directly, our framework equips the Builder with meta-skills learned from Target execution feedback. These reusable principles guide the Builder in constructing harnesses for new test tasks.
Instructions
Memory
Context
Composed tools
Execution control
Verification & recovery
Workspace
Initial H0
Generic prompt
Empty store
Default history
Native tools
Default loop
Native rules
Empty scratch
Builder B refines what?
Specific task guidance
Record format; storage and retrieval rules
History-selection rules
Tool definitions; call sequences
Execution phases; tool visibility
Submission checks; failure handling
Initial files; setup actions
Example Implement
“Draft before polishing”
Recent trial buffer
Budgeted history window
Cross-file auditor
Hide failing tools
Audit-gated submission
Templates; validator scripts
Table 1: Overview of the harness refinement space, including neutral defaults in H0 , Builder-editable mechanisms, and illustrative examples. The underlying benchmark tools, evaluation budget, and scoring rules remain fixed.
Benchmark
Tasks
Development
Test
Harness-Bench
106
11
95
NewtonBench
324
32
292
Table 2: Benchmark statistics and splits.
Method
Harness-Bench
NewtonBench
Macro Avg.
Gemini
Qwen
GPT-OSS
Gemini
Qwen
GPT-OSS
Native environment
37.35
69.63
52.46
55.48
54.11
39.38
51.40
Builder, no skills
53.29
71.51
63.02
56.16
50.34
43.84
56.36
Direct Builder skills, retrieved
46.66
68.98
59.13
58.22
48.29
33.22
52.42
Direct Builder skills, all
42.38
68.70
58.34
58.56
55.82
35.96
53.29
Independent Structured Target skills, retrieved
41.95
72.51
45.21
67.81
53.42
36.30
52.87
Table 3: Test performance (%) on every task in the fixed test splits with GPT-5.6-Sol as Builder in all settings. Bold marks the best result in each model–benchmark column.
Figure 2 : Test performance as Builder experience is refined. The two NewtonBench curves improve mainly after the second pass; Harness-Bench displays both steady improvement and late regression.
Figure 3 : (a) Same-model self-evolution compares no-skill construction, independent Target skills, and Builder meta-skills. Meta-skill bar labels show absolute scores and gains over no-skill construction. (b) Cross-Builder transfer on Harness-Bench and NewtonBench. Top: test scores without (outlined bars) and with (filled bars) Builder meta-skills. Bottom: paired gains in percentage points with 95% task-bootstrap confidence intervals; dashed lines indicate zero gain.
Target model
Full Harness
w/o Controller
95% CI
w/o Memory + Context
95% CI
Gemini-3.6-Flash
68.84
55.48-13.36
[−18.49,−8.22]
66.78-2.05
[−7.53,3.08]
Qwen-Flash
63.70
58.90-4.79
[−10.27,0.68]
59.59-4.11
[−9.25,1.03]
GPT-OSS-120B
50.68
49.66-1.03
[−6.51,4.11]
54.11+3.42
[−2.05,8.90]
Table 4: Harness component ablations on NewtonBench. We report test scores (%), changes (ablation minus full) using the main-table full-bank execution as the reference, and unadjusted 95% paired task-bootstrap confidence intervals for these changes. Intervals containing zero indicate uncertainty in the direction of the average effect.
Figure 4 : (a) Outcome transitions from no-skill to full-bank meta-skill Builders, aggregated over three Targets. Band widths represent paired-task counts; in-band numbers highlight the largest cross-outcome flows. “Zero” includes zero scores and construction failures; “Invalid” includes missing or unparsable submissions and construction failures. (b) Category-level effects across all six settings. Rows represent Harness-Bench task categories and NewtonBench physical mechanisms; columns represent Targets. Bars show score differences between full-bank and no-skill Builders with unadjusted 95% task-bootstrap confidence intervals. Purple dashed lines indicate setting’s overall effect.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Model turns
Tool calls
Tokens
Wall time
Harness-Bench
30
30
96,000
1,200 s
NewtonBench
12
10
192,000
3,600 s
Appendix
Table 5: Per-task Target limits. Each model call permits at most 16,384 output tokens; NewtonBench additionally limits one experiment call to 20 queried items. Limits are identical across conditions within a setting.