OOM-RL II: Reality Is an Oracle, Not a Debugger Provenance-Constrained Diagnosis in Continually Evolving Agent-Engineered Systems
Abstract
Reality may establish that an outcome occurred without identifying which evolving procedure produced it or why. This distinction matters in production ML systems whose code, configuration, and artifacts change while external feedback accumulates. We examine it in a human-directed, agent-engineered quantitative trading system, using oracle to mean an external source of realized outcomes rather than a complete correctness specification. Across one year, the account gained and outperformed a broad market index, while annual alpha was not statistically distinguishable from zero under the main retrospective specification. Retrospectively selected subperiods include adverse relative performance and conditional candidate-level weakness under declared approximate references. Engineering records document changes during the episode, and complete recommendation-to-runtime binding is unavailable. The archive does not establish a common frozen instance or a unique cause. The case motivates an outcome--diagnosis gap: outcome evidence, evaluated-object identity, and causal explanation support distinct claims. We distinguish frozen instances, pre-specified adaptive procedures, and ad-hoc development; organize archive-relative claim identifiability and an evidence hierarchy; and propose a prospective production-binding protocol. An illustrative compatible-history example shows how factual binding can resolve a recommendation's referent without supplying its counterfactual effect. The protocol is proposed rather than prospectively validated. External feedback constrains outcome claims, while provenance and additional identification structure determine the resolution of diagnosis.
Figures & tables
| Period | Account | Index | CAGR 252 | Sharpe | Ann. | HAC3 | |
|---|---|---|---|---|---|---|---|
| Full year | 241 | +11.1391% | -6.0997% | +11.6762% | 0.6726 | +14.4911% | 0.27526 |
| Initial | 94 | +13.1731% | +1.5075% | +39.3406% | 2.0598 | +30.0744% | 0.03339 |
| Mar–Jun | 82 | -14.7887% | +5.7058% | -38.8486% | -2.8108 | -58.1058% | 0.01778 |
| Jul–Sep | 65 | +15.2461% | -12.4876% | +73.3488% | 3.6873 | +76.5417% | 0.00230 |
| Month | Candidate scope | Median excess | Beat fraction | |
|---|---|---|---|---|
| April | all candidates | 35 | -2.1250 pp | 28.57% |
| April | primary | 12 | -2.5038 pp | 25.00% |
| June | all candidates | 45 | -1.9593 pp | 28.89% |
| June | primary | 15 | -1.9593 pp | 33.33% |
| Record date | Source identifier | Recorded change |
|---|---|---|
| 10 Mar 2026 | Git 59e2977 | Ensemble prediction/backtest alignment and normalization changes. |
| 17 Mar 2026 | Git 7af9532 | Prediction-loader refactoring toward shared Qlib Recorder infrastructure. |
| 21 Mar 2026 | Git 2194d1d | Predict-only source-artifact loading and model persistence changes. |
| 30 Mar 2026 | Git 2b5b23a | Per-mode experiment lookup for recorders. |
| 27 May 2026 | Configuration note | Hyperparameter revisions and reported production retraining. |
| 3 Jun 2026 | Git c2fa67a | Percentile-rank normalization introduced as the default fusion option. |
| Claim | Level | Status | Basis / missing requirement |
|---|---|---|---|
| Annual account outcome | 1 | Observed | Recorded returns; no persistent-alpha inference. |
| April/June T+10 weakness | 2 | Conditional | Signal-close, scopes, approximate EW; not a market-independent cause. |
| June poor under random ranking | 2 | Descriptive | 272/5,000 lower paths; not a matched null. |
| March loss proves relative failure | 2 | Unsupported | Account exceeded both cited March references. |
| Universal Buy failure; optimal Drop3 | 2 | Unsupported | Mixed horizons/scopes; non-monotonic replay. |
| One frozen instance spans incident | Identity condition | Not established | Action binding is unresolved across the archive. |
| Group | Required fields / bindings |
|---|---|
| Information timing | Signal available-after, market as-of, planned execution date, calendar identity/version. |
| Procedure identity | Procedure ID, source commit, configuration hash, prediction-loader version. |
| Model state | Ensemble identity/members, artifact IDs, training mode/cutoff, recorder IDs. |
| Decision context | Universe policy/snapshot hash, primary/alternate candidates, base and ensemble prediction hashes, final rankings. |
| Human/action path | Suggested orders, overrides/substitutions, actual orders/fills, observed constraints, fee model and actual fees with distinct roles. |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Month | Account | CSI300 | Approx. EW | |
|---|---|---|---|---|
| 2025-10 | 17 | +3.0224% | -0.0004% | +1.5938% |
| 2025-11 | 20 | -1.6414% | -2.4568% | -2.3050% |
| 2025-12 | 23 | +4.3952% | +2.2816% | +2.0311% |
| 2026-01 | 20 | +5.4351% | +1.6501% | +2.7022% |
| 2026-02 | 14 | +1.4690% | +0.0916% | +2.2296% |
| 2026-03 | 22 | -4.2922% | -5.5321% | -6.0845% |
| Baseline close | Frozen basket | Production | Frozen minus production | |
|---|---|---|---|---|
| 2 Mar 2026 | 81 | -16.6910% | -15.4779% | -1.2131 pp |
| 1 Apr 2026 | 59 | -7.2509% | -12.0480% | +4.7970 pp |
| 30 Apr 2026 | 39 | -11.9045% | -12.5792% | +0.6747 pp |
| Window | Annual | OLS | HAC(3) | HAC(5) | HAC(10) | |
|---|---|---|---|---|---|---|
| Full year | 241 | +14.4911% | 0.28244 | 0.27526 | 0.30006 | 0.32239 |
| Initial | 94 | +30.0744% | 0.09396 | 0.03339 | 0.02222 | 0.00893 |
| Mar–Jun | 82 | -58.1058% | 0.02268 | 0.01778 | 0.02319 | 0.02642 |
| Jul–Sep | 65 | +76.5417% | 0.00242 | 0.00230 | 0.00349 | 0.00473 |
| Claim | Status | Basis / limitation |
|---|---|---|
| Annual account outcome | Observed | Recorded daily account series, fixed annual period. |
| March loss proves system-specific failure | Not supported | Account outperformed both cited March references despite negative absolute return. |
| April/June T+10 candidate-relative weakness | Supported conditionally | Signal-close anchor, stated candidate scope, approximate gross EW reference. |
| Universal cross-horizon Buy failure | Not supported | Horizon results are mixed across anchors/scopes. |
| Single stock was not causal | Not established | No leave-one-name-out or concentration identification is established. |
| Sell side was irrelevant | Not established | April and June Sell summaries differ materially. |