Measuring Learned Monotone Temporal Aggregation at Matched Admissibility
Organizations: This work was conducted in a personal capacity; the views expressed are the author’s own.
Abstract
Risk regulation imposes directional constraints on scores; we adopt their strict per-input form -- the score monotone non-decreasing in every exposure input -- as a normative commitment. Deployed pipelines -- monotone hand-crafted aggregates feeding sign-constrained gradient boosting -- already satisfy it by composition, so constrained-versus-unconstrained comparisons price a guarantee the incumbent has for free. We instead hold admissibility fixed on both sides and measure what learning the aggregation is worth. Our instrument is a recurrent network whose state is classical risk statistics (an exponentially weighted moving average and a high-water mark with learned transforms), monotone by construction in every input and per MC-dropout sample. The central finding, by functional regression, is a subsumption boundary: a learned monotone channel reproduces the geometrically weighted separable family of hand-crafted statistics, one channel per member, to Spearman , approximates window statistics with measurable ceilings, and fails at consecutivity () and time localization (0.628), both structural, and at the exposure floor (0.829), a learnability boundary. One explicit admissible basis repairs each failure (rank correlation 1.000). In or near the separable family, learned and engineered aggregation are substitutes, and the learned channel is never statistically behind at full sample size and specified capacity. Its advantages are incumbent-specific: a committed grid pays up to 0.019 AUC in decay regions it leaves uncovered (the learned channel stays within 0.004 of the strongest engineered consumer at every swept point); the highest-dimensional comparator degrades fastest with scarce data; and beyond the training support, grid-fed tree-ensemble scores go flat while a strictly increasing head keeps ranking. No single incumbent is dominated on all three axes.
Figures & tables
| Regime | Logit mechanism | Targets |
|---|---|---|
| A (recency) | ||
| B (peak) | ||
| C (mixed) | EWMA and HWM jointly | |
| D (threshold stress) | ||
| E (diminishing returns) | saturating EWMA/HWM | |
| F (cumulative burden) | cumulative sum |
| Exp | Claim | Regimes | Protocol | Operators | Models |
|---|---|---|---|---|---|
| E1 | C1 | B, D | clean | rand., adv., coord., P1–P5, MC drop. | core + PGL + dropout var. |
| E2 | C2 | A (swept), C | clean | coordinate, probes | proposed, mono. aggregate |
| E3 | C3 | A–E | clean | predictive only | core + lattice |
| E4 | C4a | C | dist. shift | predictive; rand., coord. | robustness set |
| E5 | C4b | C | missingness | predictive; rand., coord. | robustness set |
| E6 | C5 | A, B, D | clean | shape recovery | proposed (learned ) |
| Claim | Experiments | Verdict |
|---|---|---|
| C1 | E1, E9 | Supported in full: zero violations on every operator and per MC-dropout sample, identically at every capacity tier. |
| C2 | E2 | Supported: inverted-U decay-sweep gap over the static aggregate, peak ( under a tanh aggregate head), closing toward both sweep ends. |
| C3 | E3, E9 | Supported: worst paired CI lower bound across the ten regime–baseline comparisons; capacity-robust from the small tier upward, with the four-parameter tiny tier’s one miss an expressiveness floor. |
| C4a | E4 | Holds as stated: guarantee exact under every ordering-preserving shift; the static aggregate is the more shift-robust scorer, and the learned channel degrades most gracefully among sequence models. |
| C4b | E5 | Supported, with its boundary quantified rather than asserted. |
| C5 | E6 | Supported in asymmetric form: and identifiable regardless of parameterization, given a non-convex monotone activation. |
| Regime B | Regime D | |||||||
| Model | Rand. | Adv. | Coord. | Probes | Rand. | Adv. | Coord. | Probes |
| Monotone exposure RNN | ||||||||
| Monotone exposure RNN (MC) | ||||||||
| Generic monotone RNN | ||||||||
| Monotone aggregate | ||||||||
| Unconstrained RNN + PGL | ||||||||
| Decay | Monotone exposure RNN | Monotone aggregate | Gap AUC |
|---|---|---|---|
| 0.10 | |||
| 0.30 | |||
| 0.50 | |||
| 0.60 | |||
| 0.70 | |||
| 0.80 |
| Regime | Mono. exp. RNN | Gen. mono. RNN | Mono. agg. | Lattice agg. | Unc. agg. | Unc. RNN | Unc. LSTM |
|---|---|---|---|---|---|---|---|
| A | |||||||
| B | |||||||
| C | |||||||
| D | |||||||
| E |
| vs. Unc. RNN | vs. Unc. LSTM | |||
|---|---|---|---|---|
| Regime | AUC | 95% CI | AUC | 95% CI |
| A | ||||
| B | ||||
| C | ||||
| D | ||||
| E | ||||
| Model | Baseline | Scale | Upper tail | Spike freq. | Mixture | Total |
|---|---|---|---|---|---|---|
| Monotone exposure RNN | 0 | 0 | 0 | 0 | 0 | 0 |
| Monotone aggregate | 0 | 0 | 0 | 0 | 0 | 0 |
| Unconstrained RNN | 1,302 | 23,758 | 15,960 | 4,083 | 6,493 | 51,596 |
| Unconstrained LSTM | 1,182 | 24,015 | 17,152 | 4,408 | 5,672 | 52,429 |
| vs. unconstrained (RNN, LSTM) | vs. monotone aggregate | ||||||
|---|---|---|---|---|---|---|---|
| Mechanism | Imputation | worst AUC | worst CI 95 hi | pass | worst AUC | worst CI 95 hi | pass |
| MCAR | zero | 10/10 | 5/5 | ||||
| MCAR | forward fill | 10/10 | 5/5 | ||||
| MCAR | mean | 10/10 | 5/5 | ||||
| Block | zero | 8/8 | 2/4 | ||||
| Block | forward fill | 8/8 | 4/4 | ||||
| , regime A ( ) | , regime B (identity) | , regime D (hinge) | , regime A | ||||
|---|---|---|---|---|---|---|---|
| Activation | corr. | RMSE | corr. | RMSE | corr. | RMSE | ( ) |
| tanh | |||||||
| sigmoid | |||||||
| softplus | |||||||
| relu | |||||||
| AUC over the 48 configurations | ||||||
|---|---|---|---|---|---|---|
| Regime | Viol. / tests | min | median | max | spread | seed std |
| A | / | |||||
| B | / | |||||
| C | / | |||||
| D | / | |||||
| Knob | Setting | A | B | C | D |
|---|---|---|---|---|---|
| identity | |||||
| monotone MLP | |||||
| identity | |||||
| hinge | |||||
| power |
| Regime | Target basis | Frozen base | Frozen target | Gap | Learned base | Residual |
|---|---|---|---|---|---|---|
| F (cumulative burden) | cumulative sum | |||||
| G (persistence) | longest run | |||||
| J (breach count) | breach count | |||||
| N (peak prominence) | HWM | |||||
| O (exposure floor) | LWM | |||||
| K (early warning) | early-window EWMA |
| Basis subset | Learned arm | Frozen arm |
|---|---|---|
| M1 (EWMA) | ||
| M2 (HWM) | ||
| M3 (EWMA+HWM) | ||
| M4 (M3 + breach count) | ||
| M5 (M4 + longest run) | ||
| M6 (M3 + cumulative + breach count) |
| Monotone exposure RNN | Unconstrained LSTM | paired non-inferiority | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Size | Regime | params | AUC | viol. | params | AUC | viol. | AUC | CI 95 | verdict |
| tiny | A | 4 | 0 | 361 | 170 | fail | ||||
| tiny | C | 4 | 0 | 361 | 1182 | pass | ||||
| small | A | 84 | 0 | 1,233 | 53 | pass | ||||
| small | C | 84 | 0 | 1,233 | 1115 | pass | ||||
| medium | A | 548 | 0 | 12,961 | 6 | pass | ||||
| G (persistence) | K (early warning) | O (exposure floor) | |
| Engineered comparators | |||
| Practice grid (constrained GBM) | |||
| Steelman grid (constrained GBM) | |||
| Raw lags (constrained GBM) | |||
| Generic monotone recurrences | |||
| Fixed-decay bank (ReLU) | |||
| Model | A | B | C | D | E | params | feat. | viol. |
|---|---|---|---|---|---|---|---|---|
| Mono. exp. RNN | 548 | – | 0 | |||||
| Mono. agg. | 241 | 5 | 0 | |||||
| Practice grid | 353 | 12 | 0 | |||||
| Steelman grid | 337 | 11 | 0 | |||||
| Practice lattice | 4,384 | 12 | 0 | |||||
| Steelman lattice | 2,312 | 11 | 0 |
| vs. Steelman GBM | vs. Raw-lags GBM | vs. Practice GBM | ||||
|---|---|---|---|---|---|---|
| Regime | AUC | 95% CI | AUC | 95% CI | AUC | 95% CI |
| A | ||||||
| B | ||||||
| C | ||||||
| D | ||||||
| E | ||||||
| Model | base | 1.5 | 2.0 | 3.0 | 5.0 | 8.0 | 2 | 3 | 5 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| Saturation rate (fraction of test scores at the maximum) | ||||||||||
| Mono. exp. RNN | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
| Mono. agg. | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
| Practice grid | 0.000 | 0.001 | 0.001 | 0.003 | 0.014 | 0.049 | 0.001 | 0.003 | 0.012 | 0.067 |
| Steelman grid | 0.000 | 0.001 | 0.002 | 0.008 | 0.033 | 0.110 | 0.002 | 0.007 | 0.028 | 0.133 |
| Practice lattice | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
| Target statistic | Channel | Spearman | |
|---|---|---|---|
| Separable accumulations | |||
| EWMA ( ) | e-channel | ||
| EWMA ( , off-grid) | e-channel | ||
| EWMA ( ) | e-channel | ||
| EWMA ( , off-grid) | e-channel | ||
| EWMA ( ) | e-channel | ||
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Range | Parameter | Range |
|---|---|---|---|
| (normal stress) | (normal log-std) | ||
| (stress normal) | (std gap) | ||
| (initial stress) | (per-step) | ||
| (AR(1) persistence) | (spike log-mean) | ||
| (normal log-mean) | (spike log-std) | ||
| (mean gap) |
| Regime | Logit |
|---|---|
| A (recency) | |
| B (peak) | |
| C (mixed) | |
| D (threshold stress) | |
| E (diminishing returns) | |
| F (cumulative burden) |
| Regime A | Regime C | |||||
|---|---|---|---|---|---|---|
| Metric | MC drop. | Ensemble | (ens MC) | MC drop. | Ensemble | (ens MC) |
| AUC | ||||||
| Log loss (NLL) | ||||||
| Brier | ||||||
| ECE | ||||||
| Uncertainty–error Spearman | ||||||
| Regime | Mono. exp. RNN | Gen. mono. RNN | Mono. agg. | Lattice agg. | Unc. agg. | Unc. RNN | Unc. LSTM |
|---|---|---|---|---|---|---|---|
| Log loss | |||||||
| A | |||||||
| B | |||||||
| C | |||||||
| D | |||||||
| E | |||||||
| E9 capacity sweep (per tier and pooled) | E5 grid | |||||
| Model | tiny | small | medium | large | total | total |
| Unconstrained RNN | 324 | 926 | 1,713 | 4,172 | 7,135 | 2,373 |
| Unconstrained LSTM | 1,352 | 1,168 | 1,188 | 794 | 4,502 | 2,425 |
| Unconstrained aggregate | 1,958 | 1,958 | 1,958 | 1,958 | 7,832 | – |