Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents
Abstract
When developers change one component of an agent, such as its controller, a learned model or its verifier, they usually judge the change by an aggregate task score. That score cannot tell whether improvement was attainable, which component lost value, or what the agent's own checks certify. We introduce a claim-specific verification audit for modular agents that plan, act, check and refine. Instead of scoring the agent, the audit scores the evidence: each conclusion is recorded with the evidence behind it, one of four verdicts (supported, unsupported, unresolved or not evaluated) and the boundary within which it holds. Three tools supply that evidence. Oracle policies measure attainable improvement under an explicitly stated action set, so that a low value can be traced to the evaluation rather than to the environment. Replacing one component at a time with a perfect counterpart locates lost value, with null results read as unresolved whenever a downstream component could mask them. A separate test asks whether the verifier's score identifies the quantity it is read as bounding. Applied to a constrained portfolio-allocation agent in a synthetic market with known hidden regimes, the audit shows that the value of perfect regime information depends on the action set used to measure it, that the scenario generator discards most of the regime signal while better local fidelity does not improve decisions, and that the runtime verifier can be bypassed with no visible change in outcomes. The contribution is the protocol and the evidential distinctions it enforces; the empirical findings are specific to the agent and environment studied.
Figures & tables
| Practice | Question it answers | Null traced to harness | Masking-aware localisation | Quality / admiss. / provenance kept apart | Verifier score identification |
|---|---|---|---|---|---|
| Aggregate benchmarking | How well does the agent do? | – | – | – | – |
| Harness auditing | Is the grader correct? | partly | – | – | – |
| Component ablation | How much does removal hurt? | – | – | – | – |
| Runtime verification | Does a run meet a spec? | – | – | partly | – |
| Decision-focused learning | Does fidelity serve decisions? | – | partly | – | – |
| Claim-specific audit | Which claims does the evidence support? | ✓ | ✓ | ✓ | ✓ |
| Verification claim | Evidence | Verdict | Boundary |
|---|---|---|---|
| Decision-relevant information is present and expressible | L2 information–action hierarchy | supported | U3, specified action set |
| ARC-Agent captures it | matched-risk frontier | unsupported | current architecture |
| Generator is an observed bottleneck | full generator bypass | supported | current trained generator |
| Improving perception would not improve downstream value | true-regime replacement | unresolved | masked by generator |
| Joint mean scenario repair improves on the all-generator baseline | coalition table | supported | U3; 6 seeds, exploratory; Sharpe only |
| Mean scenario superadditivity | Shapley–Taylor interval | unresolved | interval includes zero at 6 seeds |
| Attack | Claim surface | Effect on decision-quality evidence | Feas. audit | Solver prov. |
|---|---|---|---|---|
| A3 gate bypass | quality-verification layer | gate bypassed; metrics benign | passed | retained |
| A1 corrupted controller | untrusted configuration | degraded, within noise | passed | retained |
| A2 poisoned memory | untrusted memory | degraded, within noise | passed | retained |
| A4 adversarial controller | untrusted configuration | moved along frontier | passed | retained |
| A5 validator compromise: specified but not executed; no robustness claim. | ||||
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Verification question | Evidence | Permitted claim | |
|---|---|---|---|
| Q1 | Was improvement attainable under an action set able to express it? | L0–L3 oracle Sharpe at fixed information, widening the action set (§ 5.1 , Table 10 ); 24 seeds, paired, confirmatory | Measured headroom depends on the action representation, not on the information alone |
| Q2 | Which component replacement recovers downstream value? | Stage-wise oracle replacement with the true regime given to every arm (§ 5.2 , Table 5 ); 10–16 seeds, paired, confirmatory | A positive replacement identifies an observed bottleneck; a null one does not identify irrelevance |
| Q3 | Does improved local fidelity verify downstream improvement? | Retention, directional cosine and vector error against paired Sharpe (§ 5.3 , Table 7 ); 6 seeds, exploratory | Improved local scores do not establish downstream improvement |
| Q4 | Do joint repairs reveal structure that one-at-a-time checks miss? | Complete coalition table with Shapley and Shapley–Taylor summaries (§ 5.4 , Tables 6 , 15 ); 6 seeds, exploratory | The joint arm improves over the all-generator baseline; its conditional contribution and the interaction magnitude remain unresolved |
| Intervention | Sharpe | Turn.% | % of gap recovered |
|---|---|---|---|
| I0 baseline (scenario moments, cap ) | n/a | ||
| I1 clean regime-conditional scenarios and moments (full generator bypass) | |||
| I2 no turnover cap | |||
| L2 regime-weight reference policy | n/a |
| Channels replaced with clean regime-conditional history | Sharpe | Sharpe | CI |
|---|---|---|---|
| none (all generator; baseline) | n/a | n/a | |
| covariance only | |||
| conditional mean only | |||
| scenarios only (tail dependence) | |||
| mean covariance | |||
| scenarios covariance |
| Quantity | Definition | Reported values | Arm |
|---|---|---|---|
| Mean-gap retention | ( / bps) | deployed pipeline | |
| same, step sweep | as above | / / ( / / ) | step-count sweep |
| same, repair arm | as above | repair arm | |
| Directional similarity | repair arm | ||
| Size ratio | (overshoot) | repair arm | |
| Relative vector error | repair arm |
| Method | Sharpe | MaxDD% | Turn.% | Frontier | Excess |
| Equal weight | 0.0 | 0.74 | |||
| Minimum variance | 2.8 | 0.38 | |||
| Gaussian QP | 19.7 | 0.60 | |||
| Block bootstrap QP | 19.6 | 0.52 | |||
| Regime-cond. bootstrap QP | 19.6 | 0.53 | |||
| ARC-Agent (full) | 18.1 | 0.60 |
| Method | Sharpe | MaxDD% | Turn.% |
|---|---|---|---|
| Risk parity | |||
| Mean–variance | |||
| Hierarchical risk parity | |||
| Historical CVaR | |||
| DRO-CVaR |
| Oracle over… | Sharpe | MaxDD% | Excess | |
|---|---|---|---|---|
| L0 | (no regime knowledge) | |||
| L1 | two-point action set (EW MinVar) | |||
| L2 | regime-conditional weights | |||
| L3 | ex-post optimal weight path | |||
| ARC-Agent |
| Attack | Sharpe | MaxDD% | Infeas. | Solver record | |
| ARC-Agent | none (base) | 0 | 100% | ||
| A1 corrupted controller | 0 | 100% | |||
| A2 poisoned memory | 0 | 100% | |||
| A3 conformal-gate bypass | 0 | 100% | |||
| A4 adversarial controller (injection proxy) | 0 | 100% | |||
| weights-emitting agent (no solver) | 345 | 0% | |||
| Setting | Value |
|---|---|
| Benchmark | |
| Regimes | (bull, stagnation, crisis); searched |
| Instruments | , of which defensive (negative crisis beta) |
| Transition matrix | Eq. ( 2 ); stationary |
| Regime persistence | ; mean crisis episode days |
| Factor drift (annualised) | |
| Method | Family | CVaR-QP | Reported in |
|---|---|---|---|
| Equal weight | baseline (classical) | n/a | Tab. 8 |
| Minimum variance | baseline (classical) | n/a | Tab. 8 |
| Risk parity | baseline (classical) | n/a | Tab. 9 |
| Mean–variance | baseline (classical) | n/a | Tab. 9 |
| Hierarchical risk parity | baseline (classical) | n/a | Tab. 9 |
| Historical CVaR | baseline (scenario) | ✓ | Tab. 9 |
| Where | Quantity | Seeds | Paired | denotes | Status |
|---|---|---|---|---|---|
| Tab. 8 | Sharpe by method | 5 | no | across-seed s.d. | exploratory |
| Tab. 5 | Sharpe by intervention | 10 | yes | standard error | confirmatory |
| Tab. 6 | Sharpe, channels | 6 | yes | CI | exploratory |
| Tab. 10 | oracle Sharpe | 24 | yes | point estimates | confirmatory |
| Tab. 11 | Sharpe under attack | 5 | no | across-seed s.d. | exploratory |
| Tab. 9 | remaining baselines | 5 | no | across-seed s.d. | exploratory |
| CI | share of | ||
| A. Component credit (Shapley values; sum to ) | |||
| conditional mean | |||
| scenario tail/dependence | |||
| covariance | |||
| B. Interaction diagnostics (Shapley–Taylor indices; not added to A) | |||
| mean scenarios | |||