Agent skills that record procedural instructions are increasingly mined from execution traces rather than curated by hand. Skill-mining pipelines often use known task outcomes or feedback to guide skill construction. In production, reliable information on whether a run has succeeded may be unavailable. We study how the sampling of execution traces, access to success or failure information, and the form of the mined skills affect downstream task performance. Holding the mining pipeline fixed, we compare six combinations of mining evidence and skill forms. Mining evidence has three levels: successful trajectories only, successes and failures with their outcome labels, or the same mix with labels withheld. Skill form has two types: an ordered workflow plan, or a declarative ontology of entities, states, and policies. We evaluate the mined skills on two enterprise benchmarks, ThinkingBox-Bench and APEX-Agents. Analysis of task-level paired differences shows that the benefits of different configurations of mining evidence and skill forms depend on the enterprise domain. On ThinkingBox-Bench, paired differences show that workflows score better than ontology by 1.7 pp, Goldilocks beats success-only evidence type by 2.4 pp and Goldilocks blind simulating skills learnt without outcomes is worse by 3.1 pp. APEX-Agents shows a moderate preference for ontologies and no clear preference between evidence regimes. Within each domain, task structure related constraints drive uneven performance with mined skills. These findings motivate tailoring meta-skills to the demands of the target tasks rather than adopting a one-size-fits-all approach.
Figures & tables
Figure 1: Study design. We vary two things, shaded: the evidence the miner sees and the form of the skill it writes, using the same mining prompts in every domain. Squares show the five runs taken from each selected training task: green passed, red failed, gray outcome hidden. Success runs come from tasks that pass 8–10 of their ten no-skill runs and Goldilocks runs from tasks that pass 3–7; Goldilocks Blind uses the same runs as Goldilocks with their outcomes hidden. The three evidence types and two skill forms give the six arms run in every domain.
Benchmark
ThinkingBox-Bench
APEX-Agents
(5 domains)
(3 domains)
No skill: pass rate (%)
48.9
26.6
Workflow plan: lift over baseline
Success
+2.4
+1.1
[0.2, 4.5]
[ −0.3 , 2.5]
Goldilocks
+3.8
+0.4
Table 1: Aggregate lift of each skill-mining arm over the no-skill baseline. Domains receive equal weight within each benchmark. The shaded row gives the baseline pass rate (%); other entries are paired differences in percentage points (Equation 2 ) with 95% intervals beneath. Bold: the interval excludes zero (unadjusted). ThinkingBox-Bench intervals use a normal approximation combining domain uncertainty and treating domains as independent; APEX-Agents intervals use task bootstraps within domains. Five runs are evaluated per task and arm; APEX-Agents controls use ten historical runs. APEX-Agents grades precede judge-service corrections (Appendix A.5 ).
ThinkingBox-Bench
APEX-Agents
Split
Neobank
Consulting
Car insurance
Retail
Hotel
Investment banking
Law
Management consulting
No skill: pass rate (%)
45.0
47.5
57.1
51.0
44.0
32.6
18.6
28.5
Workflow plan
Success
+5.4
+4.9
+0.7
+1.0
−0.2
−2.2
+1.4
+4.1
[ −1.0 , 11.8]
[1.2, 8.7]
[ −3.7 , 5.0]
[ −4.1 , 6.3]
[ −4.4 , 3.8]
[ −4.3 , −0.1 ]
[ −0.8 , 3.5]
[1.2, 7.1]
Goldilocks
+7.3
+0.7
+3.7
+7.1
+0.2
−0.7
−0.6
+2.5
Table 2: Lift over the no-skill baseline within each ThinkingBox-Bench and APEX-Agents domain. Shaded rows give baseline pass rates (%); other entries give paired lift in percentage points with 95% task-bootstrap intervals beneath. Bold: the interval excludes zero (unadjusted); red: a loss. Each test task has five runs per arm; APEX-Agents controls use ten historical runs. Goldilocks Blind uses the Goldilocks traces with outcome labels removed. APEX-Agents grades precede judge-service corrections. Pass counts and inference details appear in Appendix A.5 .
Benchmark
ThinkingBox-Bench
APEX-Agents
(5 domains)
(3 domains)
Workflow plan − ontology
+1.7
−0.7
[0.3, 3.1]
[ −1.6 , 0.2]
Success − Goldilocks
−2.4
+0.4
[ −4.1 , −0.8 ]
[ −0.7 , 1.5]
Goldilocks − Goldilocks Blind
+3.1
+0.3
Table 3: Aggregate paired comparisons between skill forms and evidence conditions. Entries are first-minus-second pass-rate differences in percentage points with 95% task-bootstrap intervals beneath. Skill-form contrasts average over three evidence conditions; evidence contrasts average over two skill forms. Domains receive equal weight within each benchmark. Negative values favor ontologies in the first row and Goldilocks in the second; positive values favor labeled Goldilocks in the third. Bold: the interval excludes zero (unadjusted). Each task has five runs per arm. APEX-Agents grades precede judge-service corrections (Appendix A.5 ).
ThinkingBox-Bench
APEX-Agents
Split
Neobank
Consulting
Car insurance
Retail
Hotel
Investment banking
Law
Management consulting
Workflow plan
+2.1
+0.8
+3.5
+2.7
−0.4
−1.0
−1.0
−0.1
− ontology
[ −1.2 , 5.4]
[ −2.4 , 4.0]
[0.4, 6.7]
[ −0.2 , 5.6]
[ −3.5 , 2.6]
[ −2.5 , 0.5]
[ −2.5 , 0.4]
[ −1.8 , 1.7]
Success
−1.5
+1.7
−3.0
−4.5
−4.9
0.0
+0.6
+0.6
− Goldilocks
[ −5.3 , 2.1]
[ −1.7 , 5.0]
[ −7.2 , 1.3]
[ −7.7 , −1.4 ]
[ −9.0 , −1.0 ]
[ −1.7 , 1.8]
[ −1.2 , 2.3]
[ −1.5 , 2.7]
Goldilocks
+2.3
+2.8
+0.4
+4.8
+5.1
+1.3
−0.7
+0.4
Table 4: Paired arm comparisons within each ThinkingBox-Bench and APEX-Agents domain. Entries are first-minus-second pass-rate differences on the same test tasks, in percentage points, with 95% task-bootstrap intervals beneath. Skill-form contrasts average over three evidence conditions; evidence contrasts average over two skill forms. Negative values favor ontologies in the first row and Goldilocks in the second; positive values favor labeled Goldilocks in the third. Bold: the interval excludes zero (unadjusted). Each task has five runs per arm. APEX-Agents grades precede judge-service corrections; aggregate contrasts appear in Table 3 .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: What the two skill forms encode, drawn schematically; quoted excerpts are verbatim from skills mined in this study. (a) A workflow plan is procedural: it orders the steps of a request’s lifecycle, from discovering facts to updating the final state, and says when to stop. (b) An operational ontology is declarative: it names the entities a request involves, such as the requester and the resources they ask about, the relationships among them, such as current access, requested access, and the governing policy, and the state changes those relationships permit. The same requester can validly gain access to one resource but must be refused, with existing access preserved, for another. The ontology leaves the order of tool calls to the agent.
Inserted text
Wording
Global scope contract (workflow arms)
This is a dataset-wide mining scope. Produce one broad skill that routes by natural-language intent across every {dataset} scenario represented in the training evidence.
Task-group scope contract (ontology arms)
This is a scenario-specific mining scope for {scope_name}. Produce a narrow skill containing only workflows supported by this scope’s training evidence; exclude unrelated {dataset} scenarios.
Success evidence
The input contains five passing trajectories from each of {task_count} Success training tasks. Mine patterns repeated across tasks when the evidence permits, and preserve uncertainty for small samples.
Goldilocks evidence
The input contains three passing and two failing trajectories from each of {task_count} tasks with 3-7 passes. Contrast the paired outcomes.
Goldilocks Blind evidence
The input contains the exact same task and trajectory cohort as the Goldilocks arm: five trajectories from each of {task_count} TRAIN tasks. The only change is that all direct outcome labels are hidden. Mine only recurring observable structure, do not infer which steps caused success, and preserve contradictions and uncertainty.
Random Blind evidence
The evidence contains five distinct trajectories for each of {task_count} randomly selected TRAIN tasks. Tasks and repetitions were selected without using outcomes. Outcome labels are unavailable; do not infer hidden grades. Extract supported patterns while preserving uncertainty and contradictions.
Appendix
Table 5: Scope contracts and evidence descriptions inserted into the mining prompts, verbatim.
Figure 3: Evidence sampling. (a) Folds are built once per dataset and shared by all arms: tasks are stratified by base scenario and baseline category and spread across five folds whose sizes differ by at most one task, unless a frozen fold assignment is reused. (b) For each fold, skill scope, and evidence condition, stage 1 keeps up to ten eligible training tasks and stage 2 keeps five runs from each. Squares show one task’s ten historical runs in the order set by that condition’s hash; outlined squares are kept, and gray squares are kept runs whose outcome labels are removed before mining. The example tasks and outcomes are illustrative.
Table 6: ThinkingBox-Bench datasets, task counts, base scenarios, and the scenario groups that define ontology mining scopes.
Skill form
Evidence
Passed / runs
Pass rate (%)
Δ (pp)
95% CI (pp)
Neobank IT support (104 tasks)
No skill (baseline)
468 / 1,040
45.0
–
–
Workflow plan
Success
262 / 520
50.4
+5.4
[ −1.0 , 11.8]
Workflow plan
Goldilocks
272 / 520
52.3
+7.3
[1.4, 13.4]
Workflow plan
Goldilocks Blind
259 / 520
49.8
+4.8
[ −0.8 , 10.6]
Ontology
Success
253 / 520
48.7
+3.7
[ −0.7 , 8.1]
Appendix
Table 7: Detailed ThinkingBox-Bench results. Pass rates pool all attempts; Δ is the mean paired difference from the no-skill control in percentage points, with its 95% task-bootstrap confidence interval. Because every task has the same number of runs within an arm, Δ equals the difference in pooled pass rates. Retail and hotel booking compare with five no-skill runs per task collected for this study, and the other domains with ten earlier runs. Holm’s adjustment across the arms of each domain’s report gives these p -values: in neobank, over the six arms of its report before the Random Blind extension, 0.11 for the workflow plan mined from Goldilocks evidence and 0.44 for each other arm; in consulting, over six arms with the exact sign-flip test, 0.075 for the workflow plan mined from Success evidence, 0.17 for the Goldilocks Blind workflow plan, and 1.00 for the other four; in car insurance, over the eight arms of its report, which include the two Random Blind arms of Appendix A.7 , as the protocol specifies, with the exact sign-flip test, 0.0501 for the Goldilocks Blind workflow plan (0.046 with the paired t -test) and 1.00 for the other seven; in retail, over six arms with the paired t -test, 0.061 for the Goldilocks workflow plan and 1.00 for the other five; and in hotel, over six arms with the paired t -test, 0.25 for the ontology mined from Success evidence, 0.43 for the Goldilocks ontology, 0.79 for the Goldilocks Blind workflow plan, and 1.00 for the other three.
Domain
Task
Comparison
Passes
What differed
Skill form
Neobank
Request needing a ticket, an approval, and a pending status
Workflow vs. ontology, all three conditions
12/15 vs. 4/15
Ontology runs often stopped after opening the ticket.
Car insurance
Vehicle addition needing underwriting review
Workflow vs. ontology, Goldilocks and Goldilocks Blind
8/10 vs. 1/10
Ontology runs often left the addition or its review hold incomplete.
Car insurance
Removal of a specific vehicle
Workflow vs. ontology, Goldilocks and Goldilocks Blind
1/10 vs. 7/10
The workflow skill had no removal plan; the ontology tied the target vehicle to its policy.
Hotel
Quiet-room preference on an existing booking
Workflow vs. ontology, Success
5/5 vs. 1/5
Only one ontology run modified the booking to record the request.
Hotel
Upgrade of three rooms in one group booking
Workflow vs. ontology, Goldilocks
1/5 vs. 5/5
Workflow runs issued 14 payments across five runs; the ontology settled the group’s total as one charge.
Appendix
Table 8: Task-level reversals discussed in Section 5 , from the analysis reports. Passes count the runs of one held-out task on each side of the comparison, pooled over the evidence conditions or skill forms named; the hotel group-upgrade row under evidence conditions pools two tasks. Investment banking, law, and management consulting are APEX-Agents domains; the others are ThinkingBox-Bench datasets.
ThinkingBox-Bench
APEX-Agents
Skill form
Neobank
Consulting
Car insurance
Retail
Hotel
Investment banking
Law
Management consulting
No skill: pass rate (%)
45.0
47.5
57.1
51.0
44.0
32.6
18.6
28.5
Workflow plan
Random Blind
+1.9
+1.9
−0.7
−0.2
−1.5
−0.3
+1.3
+4.5
[ −3.1 , 6.9]
[ −1.7 , 5.5]
[ −5.7 , 4.4]
[ −5.1 , 4.7]
[ −6.7 , 3.3]
[ −2.7 , 2.0]
[ −1.3 , 3.8]
[1.9, 7.2]
Operational ontology
Appendix
Table 9: Random Blind lift over the no-skill baseline within each domain. The shaded row gives baseline pass rates (%); other entries give paired lift in percentage points with 95% task-bootstrap intervals beneath. Bold: the interval excludes zero (unadjusted); red: a loss.
Figure 4: Skill reads in all five ThinkingBox-Bench datasets and three APEX-Agents domains across the six main arms. Bars partition runs into own-scope-only, own-scope-plus-other, other-scope-only, and no skill read. Own scope is the global workflow skill or the task’s assigned ontology scenario group (ThinkingBox-Bench) or domain (APEX-Agents). Panel headers give runs per arm, pooled over all five folds with five repetitions per task. All runs have known telemetry; missing telemetry is not counted as no read. Segment widths retain exact proportions, including rare own-plus-other reads.
Dataset/domain
Success
Goldilocks
Goldilocks Blind
Judge errors
Workflow
Neobank
258/258
248/248
261/261
0/0/0
Consulting
255/255
277/277
300/300
0/0/0
Car insurance
213/213
198/198
181/181
0/0/0
Retail
235/235
205/205
243/243
0/0/0
Hotel
285/292
282/290
297/306
0/0/0
Appendix
Table 10: Recorded failures after an expected-scope skill read. Each entry gives failures with own-scope recall divided by all failures with known telemetry and usable recorded grades, pooled across five folds. Own scope is the global workflow skill or the assigned ontology scope; additional skill reads do not invalidate recall. Judge errors lists excluded APEX-Agents judge-service-error grades in Success/Goldilocks/Goldilocks Blind order. These exclusions apply only to this descriptive denominator, not headline outcomes. Read-conditioned outcomes are not causal effects.