Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations
Organizations: Department of Artificial Intelligence Wrocław University of Science and Technology
Abstract
Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against the one it was built for. Reported robustness scores, therefore, answer different questions and cannot be compared. We propose a unified cross-family evaluation protocol that holds factual instances and generated counterfactuals fixed while testing every method against the same eight types of model change. The benchmark compares six robust methods and two standard baselines on four tabular datasets. It characterizes every changed classifier through its outputs and reports empirical robustness together with coverage, base validity, and proximity. We find that relative performance and failure modes vary across change families. Bounded parameter perturbations change 0.95% of test predictions on average, compared with 4.9% for bootstrap retraining. Methods with guarantees for these perturbations do not necessarily transfer to other changes. RobX transfers most consistently in our experiments, although greater stability can require larger interventions. We argue that robust CFE methods should be evaluated through a common protocol that specifies the model changes, measures their realized behavioral magnitude, and keeps generation performance separate from robustness.
Figures & tables
| Method | Representation of model change | Robustness formulation | Changed models in original evaluation |
|---|---|---|---|
| ROAR-LIME | Bounded coefficients and intercept of a local linear surrogate | Objective under the most unfavorable allowed change | Retrained on temporally or geographically shifted data |
| RBR | Wasserstein ambiguity set around a Gaussian mixture | Objective under the most unfavorable allowed distribution | Retrained on shifted data |
| RobX | Local input stability used as a proxy for retraining | Stability threshold | Retrained after small changes to the training data |
| RNCE | Interval-bounded parameter changes | Validity certificate for every allowed change | Incremental, leave-one-out, and full retraining |
| AP S | Random bounded parameter perturbations | Probabilistic certificate | One model retrained on added data |
| BetaRCE | User-defined distribution of admissible classifiers | Probabilistic bound | Retrained with new data, hyperparameters, or architectures |
| Family | Construction |
|---|---|
| New initialization | Same data, architecture, and training configuration but different seed |
| Bootstrap | Bootstrap resampling of the training set |
| Data deletion | Random deletion of 1%, 5%, or 10% of the training observations |
| Data addition | Addition of 25%, 50%, 75%, or 100% of the reserved update pool |
| Label update | Random label flips for 1%, 5%, or 10% of the training observations |
| Training configuration | Prespecified combinations of optimizer, learning rate, and weight decay |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Split sizes | Factuals | Test balanced accuracy | ||
|---|---|---|---|---|---|
| Breast Cancer | 569 | 30 | 284/57/57/171 | 312 (60–64) | 0.967/0.973/0.988 |
| Diabetes | 768 | 8 | 384/76/77/231 | 281 (50–61) | 0.712/0.721/0.736 |
| Wine Quality | 6,497 | 11 | 3248/649/650/1950 | 1,250 (250) | 0.735/0.741/0.749 |
| HELOC | 8,291 | 20 | 4145/829/829/2488 | 1,250 (250) | 0.727/0.730/0.735 |
| Change family | Hard disagreement (%) | Probability MAE |
|---|---|---|
| Bootstrap | 4.92 1.25 | 0.039 0.010 |
| Architecture | 4.69 1.54 | 0.038 0.017 |
| New initialization | 4.57 0.93 | 0.036 0.005 |
| Label update | 3.69 1.99 | 0.045 0.025 |
| Training configuration | 3.10 1.99 | 0.028 0.019 |
| Data addition | 2.18 0.96 | 0.017 0.007 |
| Method | Mean SD | Median | 90th percentile | Timeouts |
|---|---|---|---|---|
| AP S | 3.909 4.643 | 2.612 | 7.542 | 43 |
| BetaRCE | 0.0198 0.0134 | 0.0159 | 0.0359 | 0 |
| KD-tree | 0.0037 0.0022 | 0.0042 | 0.0050 | 0 |
| RBR | 1.754 0.909 | 1.860 | 2.900 | 0 |
| RNCE | 0.262 4.200 | 0.0007 | 0.0008 | 0 |
| ROAR-LIME | 0.086 0.060 | 0.074 | 0.151 | 0 |
| Method | Coverage | Base valid | Emp. robust. | E2E robust. | Distance |
|---|---|---|---|---|---|
| Breast Cancer | |||||
| AP S | 90.4 19.8 | 100.0 0.0 | 58.5 3.6 | 53.1 12.8 | 0.689 0.201 |
| BetaRCE | 100.0 0.0 | 100.0 0.0 | 97.0 0.8 | 97.0 0.8 | 1.860 0.052 |
| KD-tree | 100.0 0.0 | 100.0 0.0 | 94.8 2.4 | 94.8 2.4 | 1.554 0.006 |
| RBR | 100.0 0.0 | 100.0 0.0 | 77.1 2.4 | 77.1 2.4 | 1.097 0.022 |
| RNCE | 100.0 0.0 | 100.0 0.0 | 98.5 0.8 | 98.5 0.8 | 1.571 0.010 |