Functional compatibility as a determinant of persistent neural learning
Organizations: School of Computing Dublin City University Dublin, Ireland
Abstract
Neural networks can acquire new capabilities while damaging existing ones, but what determines whether new learning persists remains unclear. We identify functional compatibility, the extent to which incoming learning can coexist with behaviour that must be preserved, as an experimentally manipulable causal determinant of persistence. From identical neural states, we vary compatibility while matching unrestricted learning opportunity and imposing a common retention requirement. Persistent learning increases with compatibility across independent directions, convolutional and transformer architectures, vision and text, and a ten-seed replication. Learning rules and retention constraints determine how much compatible opportunity is retained, whereas nonlinear geometry limits the matched intervention at larger update norms. Functional compatibility therefore reframes stability-plasticity from preventing forgetting to determining which new learning can coexist with existing function and persist.
Figures & tables
| System | Mean at request 0 | Nonzero MAE | Maximum error | Mean spread | Maximum spread |
|---|---|---|---|---|---|
| CNN | 0.010681 | 0.283% | 0.391% | ||
| ViT | 0.041266 | 0.282% | 0.396% | ||
| Text transformer | 0.000083 | 0.268% | 0.390% |
| CNN | ViT | Text transformer | |
|---|---|---|---|
| 0 | 0.000 [0.000,0.000] | 0.000 [0.000,0.000] | 0.000 [0.000,0.000] |
| 0.01 | 0.072 [0.008,0.136] | 0.682 [0.592,0.772] | 0.012 [0.002,0.021] |
| 0.05 | 0.999 [0.981,1.016] | 0.996 [0.987,1.005] | 0.131 [0.103,0.158] |
| 0.10 | 1.014 [1.010,1.017] | 1.001 [1.000,1.002] | 0.287 [0.233,0.341] |
| 0.25 | 1.014 [1.010,1.017] | 1.001 [1.000,1.002] | 0.728 [0.660,0.795] |
| 0.50 | 1.014 [1.010,1.017] | 1.001 [1.000,1.002] | 0.965 [0.945,0.985] |
| System | Modality | Seeds | Causal states/seed | Natural states/seed |
|---|---|---|---|---|
| CIFAR-10 CNN | vision | 11, 29, 47, 71, 101 | 50 | 50 |
| CIFAR-10 ViT | vision | 11, 29, 47, 71, 101 | 50 | 50 |
| Text transformer | text | 11, 29, 47, 71, 101 | 50 | 50 |
| System | Mean at request 0 | Nonzero MAE | Nonzero max error | Mean spread | Max spread |
|---|---|---|---|---|---|
| CIFAR-10 CNN | 0.010681 | 0.283% | 0.391% | ||
| CIFAR-10 ViT | 0.041266 | 0.282% | 0.396% | ||
| Text transformer | 0.000083 | 0.268% | 0.390% |
| Method | CIFAR-10 CNN | CIFAR-10 ViT | Text transformer | Cross-system |
|---|---|---|---|---|
| Projection / AFM-base | 1.014 [1.010, 1.017] | 1.001 [1.000, 1.002] | 0.965 [0.945, 0.985] | 0.993 [0.980, 1.006] |
| Unrestricted | 0.606 [0.591, 0.622] | 0.586 [0.581, 0.592] | 0.509 [0.488, 0.529] | 0.567 [0.542, 0.592] |
| Linearized distillation | 0.790 [0.772, 0.809] | 0.606 [0.567, 0.645] | 0.558 [0.542, 0.573] | 0.651 [0.593, 0.710] |
| EWC-prox | 0.570 [0.564, 0.577] | 0.584 [0.580, 0.589] | 0.506 [0.485, 0.528] | 0.554 [0.533, 0.574] |
| Replay | 0.013 [-0.013, 0.039] | 0.005 [-0.026, 0.036] | 0.002 [-0.003, 0.006] | 0.007 [-0.004, 0.017] |
| DER++ | 0.019 [-0.015, 0.053] | 0.010 [-0.019, 0.038] | 0.002 [-0.004, 0.008] | 0.010 [-0.001, 0.021] |
| CIFAR-10 CNN | CIFAR-10 ViT | Text transformer | |
|---|---|---|---|
| 0 | 0.000 [0.000, 0.000] | 0.000 [0.000, 0.000] | 0.000 [0.000, 0.000] |
| 0.01 | 0.072 [0.008, 0.136] | 0.682 [0.592, 0.772] | 0.012 [0.002, 0.021] |
| 0.05 | 0.999 [0.981, 1.016] | 0.996 [0.987, 1.005] | 0.131 [0.103, 0.158] |
| 0.10 | 1.014 [1.010, 1.017] | 1.001 [1.000, 1.002] | 0.287 [0.233, 0.341] |
| 0.25 | 1.014 [1.010, 1.017] | 1.001 [1.000, 1.002] | 0.728 [0.660, 0.795] |
| 0.50 | 1.014 [1.010, 1.017] | 1.001 [1.000, 1.002] | 0.965 [0.945, 0.985] |
| System | Req. | Realized | Mean | Mean | Min. margin | Accept | Margin pass | Finite |
|---|---|---|---|---|---|---|---|---|
| CIFAR-10 CNN | 0.00 | 0.010681 | 0.00524 | 0.9519 | -0.003857 | 97.2% | 78.6% | 97.2% |
| CIFAR-10 CNN | 0.10 | 0.100000 | 0.08937 | 1.0000 | 0.008058 | 100% | 100% | 99.6% |
| CIFAR-10 CNN | 0.25 | 0.250000 | 0.24836 | 1.0000 | 0.083615 | 100% | 100% | 99.6% |
| CIFAR-10 CNN | 0.50 | 0.499998 | 0.51097 | 1.0000 | 0.241686 | 100% | 100% | 99.6% |
| CIFAR-10 CNN | 0.75 | 0.749999 | 0.76107 | 1.0000 | 0.407596 | 100% | 100% | 99.6% |
| CIFAR-10 CNN | 1.00 | 0.999986 | 0.99974 | 1.0000 | 0.593211 | 100% | 100% | 99.6% |
| System | Mean | range | Mean | Median | Min. margin | |
|---|---|---|---|---|---|---|
| CIFAR-10 CNN | 250 | 0.228 | [0.067, 0.443] | 0.552 | 0.348 | 0.0995 |
| CIFAR-10 ViT | 250 | 0.422 | [0.278, 0.618] | 0.670 | 0.277 | 0.0848 |
| Text transformer | 250 | 0.950 | [0.903, 0.976] | 0.107 | 0.119 | 0.0199 |
| Method | CIFAR-10 CNN | CIFAR-10 ViT | Text transformer |
|---|---|---|---|
| AFM | 100.0% | 100.0% | 100.0% |
| Projection | 12.0% | 34.8% | 0.0% |
| Linearized distillation | 11.6% | 32.0% | 0.0% |
| EWC-prox | 0.0% | 0.0% | 0.0% |
| Replay | 0.0% | 0.0% | 0.0% |
| DER++ | 0.0% | 0.0% | 0.0% |
| System | Causal mean | Causal median | Natural mean | Natural median |
|---|---|---|---|---|
| CIFAR-10 CNN | 0.1199 | 0.1204 | ||
| CIFAR-10 ViT | 0.1582 | 0.1535 | ||
| Text transformer | 0.0270 | 0.0266 |
| Method | CNN | ViT | Text | Pooled | |
|---|---|---|---|---|---|
| Projection / AFM-base | 0.00 | 0.000 | 0.000 | -0.000 | -0.000 |
| Projection / AFM-base | 0.01 | 0.072 | 0.682 | 0.012 | 0.255 |
| Projection / AFM-base | 0.05 | 0.999 | 0.996 | 0.131 | 0.708 |
| Projection / AFM-base | 0.10 | 1.014 | 1.001 | 0.287 | 0.767 |
| Projection / AFM-base | 0.25 | 1.014 | 1.001 | 0.728 | 0.914 |
| Projection / AFM-base | 0.50 | 1.014 | 1.001 | 0.965 | 0.993 |
| System | Method | Mean | Median | Retention pass |
|---|---|---|---|---|
| CIFAR-10 CNN | AFM | 0.937 | 0.348 | 100.0% |
| CIFAR-10 CNN | Projection | 1.787 | 0.631 | 12.0% |
| CIFAR-10 CNN | Linearized distillation | 1.789 | 0.632 | 11.6% |
| CIFAR-10 CNN | EWC-prox | 1.075 | 1.008 | 0.0% |
| CIFAR-10 CNN | Replay | -13.393 | -3.674 | 0.0% |
| CIFAR-10 CNN | DER++ | -2.477 | -0.223 | 0.0% |
| CNN /slope | ViT /slope | Text /slope | |
|---|---|---|---|
| 0.01 | 0.3718 / 0.0725 | 0.2467 / 0.2740 | 0.4353 / |
| 0.05 | 0.0922 / 0.6876 | 0.0409 / 0.8535 | 0.5321 / 0.0927 |
| 0.10 | 0.0577 / 0.8917 | 0.0220 / 0.9566 | 0.3445 / 0.3367 |
| 0.25 | 0.0400 / 1.0130 | 0.0254 / 1.0133 | 0.1151 / 0.8285 |
| 0.50 | 0.0379 / 1.0465 | 0.0206 / 1.0294 | 0.0302 / 0.9861 |
| 1.00 | 0.0347 / 1.0556 | 0.0205 / 1.0295 | 0.0178 / 1.0033 |
| Equal-system pooled slope | 95% CI | |
|---|---|---|
| 0.01 | 0.36331 | [0.35075,0.37667] |
| 0.05 | 0.77171 | [0.76572,0.77759] |
| 0.10 | 0.82114 | [0.81433,0.82756] |
| 0.25 | 0.93121 | [0.92425,0.93841] |
| 0.50 | 0.99360 | [0.99123,0.99600] |
| 1.00 | 1.00379 | [1.00326,1.00442] |
| Dataset | AFM | CCL-DC | MKD | FGH | aL-SAR | LPR | OCAR |
|---|---|---|---|---|---|---|---|
| CORe50 | 0.2478 | 0.2016 | 0.1799 | 0.2546 | 0.1801 | 0.1640 | 0.1436 |
| CLEAR-10 | 0.3807 | 0.3731 | 0.3513 | 0.3773 | 0.3572 | 0.3485 | 0.3164 |
| CLAD-C | 0.5129 | 0.4921 | 0.4931 | 0.3058 | 0.3296 | 0.4694 | 0.3747 |
| System | Negative requested-zero empirical margins |
|---|---|
| CNN | 52 |
| ViT | 1 |
| Text transformer | 12 |
| Total | 65 |
| System | Natural-norm fraction | Feasible / attempted | Seeds with feasibility | Target norm |
|---|---|---|---|---|
| CNN | 0.01 | 1/500 (0.2%) | 1 | 0.00060570 |
| CNN | 0.10-1.00 | 0/1500 | 0 | 0.00605702-0.06057018 |
| ViT | 0.01 | 231/500 (46.2%) | 10 | 0.00039144 |
| ViT | 0.10-1.00 | 0/1500 | 0 | 0.00391441-0.03914410 |
| System | Matched slope | 95% CI |
|---|---|---|
| CIFAR-10 CNN | 1.01377 | [1.01027, 1.01727] |
| CIFAR-10 ViT | 1.00082 | [0.99999, 1.00165] |
| Text transformer | 0.96497 | [0.94527, 0.98466] |
| Cross-system descriptive mean | 0.99319 | [0.98043, 1.00594] |
| System | ||||
|---|---|---|---|---|
| CIFAR-10 CNN | 1.01290 | 1.01032 | 1.00359 | 0.99390 |
| CIFAR-10 ViT | 1.00135 | 0.99886 | 0.99567 | 0.98347 |
| System | Mean causal norm | Median natural norm | Causal/natural ratio |
|---|---|---|---|
| CIFAR-10 CNN | 0.0605702 | ||
| CIFAR-10 ViT | 0.0391441 | ||
| Text transformer | 0.0146534 | 0.0741306 |
| CIFAR-10 CNN | CIFAR-10 ViT | Text transformer | |
|---|---|---|---|
| 0.01 | 0.3718 / 0.0725 | 0.2467 / 0.2740 | 0.4353 / |
| 0.05 | 0.0922 / 0.6876 | 0.0409 / 0.8535 | 0.5321 / 0.0927 |
| 0.10 | 0.0577 / 0.8917 | 0.0220 / 0.9566 | 0.3445 / 0.3367 |
| 0.25 | 0.0400 / 1.0130 | 0.0254 / 1.0133 | 0.1151 / 0.8285 |
| 0.50 | 0.0379 / 1.0465 | 0.0206 / 1.0294 | 0.0302 / 0.9861 |
| 1.00 | 0.0347 / 1.0556 | 0.0205 / 1.0295 | 0.0178 / 1.0033 |
| System | Mean matched slope | 95% CI |
|---|---|---|
| CIFAR-10 CNN | 1.01544 | [1.01324, 1.01778] |
| CIFAR-10 ViT | 1.00064 | [1.00010, 1.00114] |
| Stronger CIFAR-10 ViT | 0.99883 | [0.99824, 0.99942] |
| Text transformer | 0.95950 | [0.95049, 0.96863] |
| Pooled slope | 95% CI | |
|---|---|---|
| 0.01 | 0.36331 | [0.35075, 0.37667] |
| 0.05 | 0.77171 | [0.76572, 0.77759] |
| 0.10 | 0.82114 | [0.81433, 0.82756] |
| 0.25 | 0.93121 | [0.92425, 0.93841] |
| 0.50 | 0.99360 | [0.99123, 0.99600] |
| 1.00 | 1.00379 | [1.00326, 1.00442] |
| System | Count |
|---|---|
| CIFAR-10 CNN | 52 |
| CIFAR-10 ViT | 1 |
| Text transformer | 12 |
| Total | 65 |
| System | Natural-norm fraction | Target norm | Feasible/attempted | Seeds with any feasible |
|---|---|---|---|---|
| CNN | 0.01 | 0.00060570 | 1/500 (0.2%) | 1 |
| CNN | 0.10 | 0.00605702 | 0/500 | 0 |
| CNN | 0.50 | 0.03028509 | 0/500 | 0 |
| CNN | 1.00 | 0.06057018 | 0/500 | 0 |
| ViT | 0.01 | 0.00039144 | 231/500 (46.2%) | 10 |
| ViT | 0.10 | 0.00391441 | 0/500 | 0 |
| Method | slope [95% CI] | slope [95% CI] |
|---|---|---|
| EWC-prox | [3.5585, 5.1115] | [ , ] |
| Linearized distillation | [2.54328, 2.65080] | [ , ] |
| Projection / AFM base | [3.03053, 3.09322] | [ , 0.455] |
| Unrestricted branch | [3.55718, 5.12805] | [ , ] |
| -CV band | States | slope | slope |
|---|---|---|---|
| Low | 78 | 2.2463 | |
| Medium | 77 | 3.1477 | |
| High | 76 |
| CORe50 | CLEAR-10 | CLAD-C | |
|---|---|---|---|
| Dataset structure | Official CORe50 stream Lomonaco and Maltoni (2017) ; classes and predeclared episodes covering new contexts, recurrence, long-dormancy return, an explicit target conflict, gradual visual drift, and capacity pressure. | Official CLEAR-10 imagery Lin et al. (2021) ; labels including background, chronological supervised buckets, and natural temporal drift. | CLAD-C chronological object classification Verwimp et al. (2023) on labeled SODA10M Han et al. (2021) ; classes and official chronological training segments. Only labeled object crops are used: no SODA10M unlabeled pretraining, no CLAD-D, and no external pretrained backbone. |
| Learner stream | Batch size ; task/session/episode metadata withheld; evaluator-only semantic regimes and validity intervals. | learner examples, per bucket, batch size , updates; bucket and period metadata withheld. | official object crops in six preserved segments with item counts 5157, 1154, 6742, 2560, 4517 and 2119; maximum batch size with partial boundary batches retained. Segment metadata and original labels remain evaluator-only. |
| Representation prefix | batches / images, then backbone freeze. | Complete first supervised episode: batches / images, then backbone freeze. Optional unlabeled bucket- pretraining is not used. | Complete first official training segment: batches / crops, then backbone freeze. |
| Candidate fitting | routed examples, fixed epochs, validation horizon , risk threshold . | routed examples, fixed epochs, validation horizon , risk threshold . | routed examples, fixed epochs, validation horizon , risk threshold . |
| Evaluation | Checkpoint matrix over active semantic regimes; held-out test and complete original-semantics evaluator. | Checkpoint at every bucket boundary; validation and held-out test examples from supervised buckets - ; next-bucket evaluation for near-future accuracy. | Checkpoint after every official training segment; the official data loader yields validation and held-out test object crops. |
| Matrix | variants per seed and seeds: jobs per dataset, jobs total across the three frozen matrices. | ||
| Dataset | Comparator | AFM HV | Comparator HV | Mean | 95% CI | Wins | Relative |
|---|---|---|---|---|---|---|---|
| CORe50 | No protection | 0.2478 | 0.2388 | +0.0090 | [0.0059, 0.0126] | 5/5 | 3.76% |
| CLEAR-10 | No protection | 0.3807 | 0.3740 | +0.0068 | [0.0053, 0.0082] | 5/5 | 1.81% |
| CLAD-C | A-GEM | 0.5129 | 0.5019 | +0.0110 | [0.0054, 0.0166] | 5/5 | 2.20% |
| Family | CORe50 | CLEAR-10 | CLAD-C |
|---|---|---|---|
| AFM | 0.2478 | 0.3807 | 0.5129 |
| No protection | 0.2388 | 0.3740 | 0.4872 |
| Matched SGD | 0.2031 | 0.3596 | 0.4714 |
| Replay | 0.1839 | 0.3503 | 0.4797 |
| A-GEM | 0.1845 | 0.3521 | 0.5019 |
| Online EWC | 0.1829 | 0.3511 | 0.4834 |
| Dataset | Point | Fgt. | BWT | Test | Online | NF | ||
|---|---|---|---|---|---|---|---|---|
| CORe50 | AFM 0.10 | 0.9648 | 0.2088 | 0.0352 | -0.0186 | 0.2291 | 0.2125 | - |
| CORe50 | AFM 0.50 | 0.9185 | 0.2458 | 0.0815 | -0.0658 | 0.2362 | 0.2594 | - |
| CORe50 | AFM 1.00 | 0.9020 | 0.2595 | 0.0980 | -0.0842 | 0.2398 | 0.2764 | - |
| CORe50 | No protection | 0.8866 | 0.2693 | 0.1134 | -0.0991 | 0.2341 | 0.2866 | - |
| CLEAR-10 | AFM 0.10 | 0.9933 | 0.3750 | 0.0067 | 0.0056 | 0.3412 | 0.3753 | 0.3263 |
| CLEAR-10 | AFM 0.50 | 0.9828 | 0.3793 | 0.0172 | -0.0013 | 0.3417 | 0.3799 | 0.3309 |
| Dataset | AFM HV | CCL-DC HV | MKD HV | FGH HV | aL-SAR HV | LPR HV | OCAR HV |
|---|---|---|---|---|---|---|---|
| CORe50 | 0.2478 | 0.2016 | 0.1799 | 0.2546 | 0.1801 | 0.1640 | 0.1436 |
| CLEAR-10 | 0.3807 | 0.3731 | 0.3513 | 0.3773 | 0.3572 | 0.3485 | 0.3164 |
| CLAD-C | 0.5129 | 0.4921 | 0.4931 | 0.3058 | 0.3296 | 0.4694 | 0.3747 |
| Method | Fgt. | BWT | Online | Test | Primary | ||
|---|---|---|---|---|---|---|---|
| AFM | 0.9127 | 0.4031 | 0.0873 | -0.0312 | 0.4092 | 0.3849 | 0.3805 |
| CCL-DC | 0.9424 | 0.3850 | 0.0576 | 0.0081 | 0.3955 | 0.3950 | 0.3556 |
| MKD | 0.9634 | 0.3604 | 0.0366 | 0.0175 | 0.3714 | 0.3916 | 0.3414 |
| LPR | 0.9538 | 0.3512 | 0.0462 | 0.0202 | 0.3634 | 0.3799 | 0.3273 |
| FGH | 0.8716 | 0.3592 | 0.1284 | -0.0215 | 0.3816 | 0.3591 | 0.3126 |
| aL-SAR | 0.9625 | 0.3018 | 0.0375 | 0.0343 | 0.3277 | 0.2934 | 0.2890 |
| Comparison | Fgt. | BWT | Online | Test | Primary | ||
|---|---|---|---|---|---|---|---|
| AFM vs CCL-DC | -3.14% | +4.72% | +2.96 pp | -3.93 pp | +3.46% | -2.55% | +7.01% |
| AFM vs MKD | -5.26% | +11.86% | +5.07 pp | -4.87 pp | +10.19% | -1.70% | +11.44% |
| AFM vs LPR | -4.31% | +14.80% | +4.11 pp | -5.14 pp | +12.62% | +1.32% | +16.26% |
| AFM vs FGH | +4.72% | +12.24% | -4.12 pp | -0.97 pp | +7.22% | +7.19% | +21.73% |
| AFM vs aL-SAR | -5.17% | +33.58% | +4.98 pp | -6.55 pp | +24.88% | +31.20% | +31.67% |
| AFM vs OCAR | -8.29% | +44.00% | +8.25 pp | -3.62 pp | +34.54% | +37.93% | +36.76% |
| Ablation | Full AFM | Ablation | Mean | 95% CI | Wins | Relative gain |
|---|---|---|---|---|---|---|
| Base-only, finite normalized budget | 0.2257 | 0.1903 | +0.0354 | [+0.0276, +0.0445] | 5/5 | +18.61% |
| Base-only, fixed-absolute budget | 0.2257 | 0.1903 | +0.0354 | [+0.0266, +0.0446] | 5/5 | +18.59% |
| Single-timescale AFM | 0.2257 | 0.2257 | +0.0000 | [-0.0008, +0.0005] | 4/5 | +0.01% |
| Dataset | Runs | Certs | Commits | Protected nonzero | Exact restore accepted/attempts | Max endpoint error | Min deployed ratio | Min analytic margin |
|---|---|---|---|---|---|---|---|---|
| CORe50 | 15 | 90 | 90 | 8621 | 8636/8829 | 0.99999494 | ||
| CLEAR-10 | 15 | 99 | 99 | 10416 | 10431/10494 | 0.99999718 | ||
| CLAD-C | 15 | 488 | 488 | 22917 | 22932/24980 | 0.99993644 |
| Component | CORe50 | CLEAR-10 | CLAD-C |
|---|---|---|---|
| Ordinary learning | 0.36% | 0.44% | 0.24% |
| Same-state no-protection comparator | 0.52% | 0.27% | 0.25% |
| Protection geometry | 37.33% | 16.12% | 42.06% |
| Candidate machinery | 1.94% | 2.17% | 2.00% |
| Finite counterfactual completion | 25.95% | 22.90% | 28.55% |
| Endpoint verification | 31.54% | 52.68% | 25.17% |
| Seed | CORe50 AFM | CORe50 no-prot. | CLEAR AFM | CLEAR no-prot. | CLAD AFM | CLAD A-GEM | |||
|---|---|---|---|---|---|---|---|---|---|
| 11 | 0.2481 | 0.2403 | +0.0078 | 0.3852 | 0.3763 | +0.0089 | 0.5065 | 0.4961 | +0.0104 |
| 29 | 0.2281 | 0.2205 | +0.0076 | 0.3693 | 0.3652 | +0.0041 | 0.5001 | 0.4951 | +0.0051 |
| 47 | 0.2520 | 0.2371 | +0.0149 | 0.3836 | 0.3778 | +0.0058 | 0.5231 | 0.5198 | +0.0033 |
| 71 | 0.2621 | 0.2586 | +0.0036 | 0.3752 | 0.3684 | +0.0069 | 0.5216 | 0.5010 | +0.0207 |
| 101 | 0.2489 | 0.2377 | +0.0111 | 0.3903 | 0.3822 | +0.0081 | 0.5133 | 0.4977 | +0.0156 |
| Dataset | Coordinate | Forgetting reduction | Relative reduction | Test difference |
|---|---|---|---|---|
| CORe50 | 0.10 | 0.0783 | 69.00% | -0.499 |
| CORe50 | 0.50 | 0.0319 | 28.12% | +0.205 |
| CORe50 | 1.00 | 0.0154 | 13.59% | +0.565 |
| CLEAR-10 | 0.10 | 0.0177 | 72.66% | +0.358 |
| CLEAR-10 | 0.50 | 0.0072 | 29.42% | +0.408 |
| CLEAR-10 | 1.00 | 0.0012 | 5.06% | +0.101 |
| Method | Persistent auxiliary state | Optimization | Principal frozen settings |
|---|---|---|---|
| OCAR | Replay memory, 128 examples | Shared learning rate | Curvature/Fisher, damping, and replay rules retained |
| LPR | Replay memory, 128 examples | Learning rate | , , preconditioner refresh every 100 iterations |
| aL-SAR | Replay memory, 4000 examples | Adam, | Update batch 16, online factor , unfreeze rate , temperature , , warm-up 50 |
| FGH | Prototype state | Adam, | Hypergradient rate 1, gradient-weight clamp 1000, one online epoch, prototype coefficient 1 |
| CCL-DC | Second learner, replay memory 1000 | AdamW, , weight decay | Replay batch 64, one memory iteration, distillation weight 2, temperature 4 |
| MKD | EMA teacher, replay memory 1000 | Adam, | Replay batch 64, distillation weight 5.5, temperature 4, EMA , correction interval 10 |
| Method | Equal-weighted three-benchmark primary score |
|---|---|
| AFM | 0.3805 |
| CCL-DC | 0.3556 |
| MKD | 0.3414 |
| LPR | 0.3273 |
| FGH | 0.3126 |
| aL-SAR | 0.2890 |
| Dataset | Method | Fgt. | BWT | Test | Online | ||
|---|---|---|---|---|---|---|---|
| CORe50 | AFM 0.10 | 0.9648 | 0.2088 | 0.0352 | -0.0186 | 0.2291 | 0.2125 |
| AFM 0.50 | 0.9185 | 0.2458 | 0.0815 | -0.0658 | 0.2362 | 0.2594 | |
| AFM 1.00 | 0.9020 | 0.2595 | 0.0980 | -0.0842 | 0.2398 | 0.2764 | |
| CCL-DC | 0.9756 | 0.2066 | 0.0244 | 0.0339 | 0.2648 | 0.2240 | |
| MKD | 0.9952 | 0.1808 | 0.0048 | 0.0374 | 0.2415 | 0.1929 | |
| FGH | 0.8755 | 0.2910 | 0.1245 | -0.1088 | 0.2493 | 0.3087 |
| Quantity | AFM | No protection |
|---|---|---|
| Pedestrian peak accuracy at boundary 2 | 77.24% | 60.34% |
| Pedestrian accuracy at boundary 3 | 0.00% | 0.00% |
| Final Pedestrian accuracy | 1.58% | 1.76% |
| Pedestrian peak-to-final forgetting | 75.66% | 58.59% |
| Mean Pedestrian margin at boundary 2 | +0.482 | +0.053 |
| Mean Pedestrian margin at boundary 3 | -4.202 | -5.587 |
| Method | Fgt. | BWT | Online | Test | |||
|---|---|---|---|---|---|---|---|
| Full AFM, | 0.9185 | 0.2458 | 0.0815 | -0.0658 | 0.2594 | 0.2362 | 0.2257 |
| Base-only, finite normalized budget | 0.9984 | 0.1906 | 0.0016 | -0.0005 | 0.1885 | 0.2007 | 0.1903 |
| Base-only, fixed-absolute budget | 0.9977 | 0.1908 | 0.0023 | -0.0004 | 0.1886 | 0.2016 | 0.1903 |
| Single-timescale AFM | 0.9205 | 0.2452 | 0.0795 | -0.0638 | 0.2589 | 0.2354 | 0.2257 |
| Operation | Median calls | Median host-observed time/call | Host-timing share |
|---|---|---|---|
| Protected-projector construction | 1,711 | 2861.552 ms | 37.89% |
| Compact-cardinal shield construction | 1,526 | 1611.751 ms | 20.58% |
| Safe-base backtracking | 1,665 | 613.933 ms | 8.60% |
| Finite-address preparation | 1,665 | 601.503 ms | 8.08% |
| Protected prestate evaluation | 1,711 | 599.454 ms | 8.06% |
| Protected poststate evaluation | 1,664 | 599.695 ms | 8.05% |
| Scenario | Seeds passed | Structural events | Observed behavior |
|---|---|---|---|
| Favourable recurrence | 3/3 | commits 6/6/6; protected steps 20/13/10 | recurrent protected learning completed |
| Semantic conflict | 3/3 | reopenings 2/3/1; splits 0/0/0 | conflict handled without an unsupported route split |
| Observable shift | 3/3 | commits 3/3/3; protected steps 41/27/27 | observable-shift control completed |
| Within-route observable shift | 3/3 | splits 1/1/1 | positive route splitting exercised on every seed |
| Capacity pressure | 3/3 | commits 11/11/11; protected steps 30/18/21 | bounded-capacity control completed |
| Identical-observation impossibility | 3/3 | commits 3/3/3; protected steps 12/8/11 | indistinguishable observations did not induce a false semantic resolution |
| Method | |||
|---|---|---|---|
| Full AFM, | 0.9185 | 0.2458 | 0.2257 |
| Base-only, normalized budget | 0.9984 | 0.1906 | 0.1903 |
| Base-only, fixed absolute budget | 0.9977 | 0.1908 | 0.1903 |
| Single-timescale AFM | 0.9205 | 0.2452 | 0.2257 |