The Harness as the Only Mutable Surface: Compliance-Bounded Self-Evolution of LLM Agents in Credit Pipelines, with a Measured Admission Gate
Organizations: Digital Economy Lab
Abstract
Self-improving LLM agents can adapt a credit pipeline to a changed rule, but an agent that rewrites itself destroys the artefact a supervisor reviews: a named change, a recorded test, an approval. We argue that self-evolution is reviewable only if it is confined to the runtime harness (instruction text, tool-call logic and primitive composition) while model weights stay fixed, so that every adaptation is a diff with a cause and a test attached. We give a dual-loop engine built on that bound, with one admission gate that writes a hash-chained record before deployment, and we measure the gate in simulation, with a simulated agent and a seeded-search proposer rather than language models. Across three families of supervisory re-interpretation at three severities, 10 seeds each, the gated loop admitted 144 of 7,449 candidate changes, none of which worsened error on held-out history, and restored the false-positive rate to the oracle level without raising missed flags in every low- and mid-severity cell. With the gate replaced by the check an unbounded system applies (fewer errors visible in recent traces), the same loops admitted 309 harmful changes and left missed flags above 10% in 49 of 90 runs: false positives fell because the screen was loosened. Evaluated on pre-shift labels, the gate rejected every candidate, so a re-interpretation must be encoded as a rule that relabels history. Parametric and scope shifts were repaired locally, a structural one only by primitive replacement; at the highest structural severity the gate's fixed tolerance blocked the correct replacement in half the seeds. We map the mechanisms to the EU AI Act's provisions for high-risk credit scoring and note that the April 2026 US model-risk guidance excludes agentic AI from its scope.
Figures & tables
| System | What evolves | Gate | Record | Gate measured |
|---|---|---|---|---|
| Reflexion [ 9 ] | memory | – | – | – |
| PromptBreeder [ 8 ] | prompts | – | – | – |
| GEPA [ 7 ] | prompts | – | – | |
| Self-Harness [ 5 ] | harness | – | – | |
| HarnessX [ 6 ] | primitives, weights | – | – | |
| EvoAgentX [ 13 ] | workflows | – | – |
| Parameter | Value |
|---|---|
| Stream | 23,000 screening alerts; shift at ; adaptation on ; held-out evaluation on |
| Retrospective pool | pre- stream; even-indexed cases (gate), odd-indexed (audit only) |
| Window / cycle | 500 cases; flags reviewed 100%; clears audited 5% |
| Execution noise | outcome flipped with probability 0.02 per case, draw shared across harness versions |
| Local loop | candidates per cluster, clusters per cycle; for a false-positive cluster: remove its hit type from scope, or ; for a miss cluster: add its hit type, or |
| Escalation | consecutive cycles without admission for a cluster |
| Arm | FP | Missed flag | Regr. | Excess err. | Recov. win. | Gate: prop. / adm. / false |
| Thr shift, mid severity, 10 seeds | ||||||
| Healthy (no shift) | 2.0 [1.6, 2.3] | 1.9 [1.6, 2.1] | 0.0 | – | – | – |
| Degraded (no adaptation) | 28.0 [27.3, 28.7] | 1.9 [1.6, 2.2] | 0.0 | 498 | 7.0 | – |
| Oracle update | 1.9 [1.6, 2.2] | 1.9 [1.6, 2.2] | 0.3 | 0 | 0.0 | – |
| Manual update (lag 2,000) | 1.9 [1.6, 2.2] | 1.9 [1.6, 2.2] | 0.3 | 295 | 4.0 | – |
| No gate (trace check only) | 1.9 [1.6, 2.2] | 1.9 {1.2–2.5} | 0.3 | 139 | 2.5 | 784 / 30 / 6 |
| Family | Severity | clear / flag | Degraded | No gate | Global only | Local only | Dual | Oracle |
|---|---|---|---|---|---|---|---|---|
| Thr | low | 1,600 / 1,850 | 13.1 / 1.9 | 2.0 / 1.9 | 13.1 / 1.9 | 2.0 / 1.9 | 2.0 / 1.9 | 2.0 / 1.9 |
| Thr | mid | 1,940 / 1,510 | 28.0 / 1.9 | 1.9 / 1.9 | 28.0 / 1.9 | 1.9 / 1.9 | 1.9 / 1.9 | 1.9 / 1.9 |
| Thr | high | 2,261 / 1,189 | 38.0 / 2.0 | 1.9 / 3.1 | 38.0 / 2.0 | 1.9 / 2.0 | 1.9 / 2.0 | 1.9 / 2.0 |
| Scope | low | 1,815 / 1,635 | 23.1 / 1.8 | 2.0 / 69.2 | 23.1 / 1.8 | 2.0 / 1.8 | 2.0 / 1.8 | 2.0 / 1.8 |
| Scope | mid | 2,223 / 1,227 | 36.9 / 1.8 | 2.0 / 79.7 | 36.9 / 1.8 | 3.7 / 1.8 | 2.0 / 1.8 | 2.0 / 1.8 |
| Scope | high | 2,627 / 823 | 46.3 / 1.9 | 1.9 / 78.8 | 46.3 / 1.9 | 1.9 / 1.9 | 1.9 / 1.9 | 1.9 / 1.9 |
| Arm | Prop. | Adm. | Rej.: no gain | Rej.: regr. | False adm. | Oracle rej. | Miss 10% | Max adm. |
|---|---|---|---|---|---|---|---|---|
| Dual loop (gated) | 7449 | 144 | 5640 | 1665 | 0 | 25 | 0 | 4 |
| Local loop only | 5315 | 111 | 2745 | 2459 | 0 | 2 | 0 | 3 |
| Global loop only | 2443 | 25 | 500 | 1918 | 0 | 28 | 0 | 1 |
| Gate on stale labels | 8490 | 0 | 8490 | 0 | 0 | 470 | 0 | 0 |
| No gate (trace check only) | 5353 | 550 | 4803 † | – | 309 | 0 | 49 | 13 |
| Severity | FP % | Missed flag % | Recovered | Oracle rej. | False adm. | |
|---|---|---|---|---|---|---|
| low | 0.005 | 2.0 [1.7, 2.3] | 1.9 [1.6, 2.2] | 10/10 | 0 | 0 |
| low | 0.010 | 2.0 [1.7, 2.3] | 1.9 [1.6, 2.2] | 10/10 | 0 | 0 |
| mid | 0.005 | 2.0 [1.7, 2.3] | 1.9 [1.6, 2.2] | 10/10 | 0 | 0 |
| mid | 0.010 | 2.0 [1.7, 2.3] | 1.9 [1.6, 2.2] | 10/10 | 0 | 0 |
| high | 0.005 | 18.1 [6.0, 30.2] | 1.8 [1.5, 2.1] | 5/10 | 24 | 0 |
| high | 0.010 | 2.1 [1.8, 2.3] | 1.8 [1.5, 2.1] | 10/10 | 0 | 0 |
| Obligation (EU AI Act) | Mechanism | Artefact handed to an examiner | Evidence here |
|---|---|---|---|
| Predetermined changes, Art. 43(4); substantial modification, Art. 3(23) | Harness-only evolution over a fixed edit space and library | Edit space, library and Eq. ( 1 ) in technical documentation | Implemented; bound enforced in code |
| Record-keeping, Art. 12 | Gate writes before deployment; hash chain | Append-only log of admissions and rejections | Implemented; 90 logs verified |
| Accuracy and robustness; feedback loops, Art. 15(1),(4) | Non-regression on ; held-out ; missed-flag rate tracked | Per-candidate test result | Measured: 0 false admissions (a weak test; see Section 10 ) |
| Traceability of the model in force | Weights pinned by digest | Model digest per decision | Design only: no model in the simulation |
| Change management in the QMS, Art. 17(1)(a) | Gate as the only path to deployment | Rejected-candidate records | Measured: rejections by reason |
| Human oversight, Art. 14 | Rule authored by the compliance function; escalation to people | Rule version per decision; escalation records | Design only |
| Metric (%) | Blind | Judge | Blind, seeds | Judge, seeds | |
|---|---|---|---|---|---|
| Silent decision | 55.1 [53.5, 56.7] | 6.5 [5.8, 7.4] | 3,582 | 54.3–56.1 | 6.5–7.4 |
| Compliance FP | 22.5 [21.8, 23.2] | 5.0 [4.6, 5.3] | 14,948 | 22.2–22.7 | 4.9–5.1 |
| Missed flag | 15.3 [14.3, 16.4] | 3.5 [3.0, 4.0] | 4,470 | 14.6–15.3 | 3.1–3.5 |
| Hallucination | 2.6 [2.4, 2.8] | 0.6 [0.5, 0.7] | 23,000 | 2.5–2.8 | 0.4–0.7 |
| Provenance compl. | 93.4 [93.1, 93.7] | 99.3 [99.2, 99.4] | 23,000 | 93.3–93.8 | 99.3–99.4 |