Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against the one it was built for. Reported robustness scores, therefore, answer different questions and cannot be compared. We propose a unified cross-family evaluation protocol that holds factual instances and generated counterfactuals fixed while testing every method against the same eight types of model change. The benchmark compares six robust methods and two standard baselines on four tabular datasets. It characterizes every changed classifier through its outputs and reports empirical robustness together with coverage, base validity, and proximity. We find that relative performance and failure modes vary across change families. Bounded parameter perturbations change 0.95% of test predictions on average, compared with 4.9% for bootstrap retraining. Methods with guarantees for these perturbations do not necessarily transfer to other changes. RobX transfers most consistently in our experiments, although greater stability can require larger interventions. We argue that robust CFE methods should be evaluated through a common protocol that specifies the model changes, measures their realized behavioral magnitude, and keeps generation performance separate from robustness.
Figures & tables
Method
Representation of model change
Robustness formulation
Changed models in original evaluation
ROAR-LIME
Bounded coefficients and intercept of a local linear surrogate
Objective under the most unfavorable allowed change
Retrained on temporally or geographically shifted data
RBR
Wasserstein ambiguity set around a Gaussian mixture
Objective under the most unfavorable allowed distribution
Retrained on shifted data
RobX
Local input stability used as a proxy for retraining
Stability threshold
Retrained after small changes to the training data
RNCE
Interval-bounded parameter changes
Validity certificate for every allowed change
Incremental, leave-one-out, and full retraining
AP Δ S
Random bounded parameter perturbations
Probabilistic certificate
One model retrained on added data
BetaRCE
User-defined distribution of admissible classifiers
Probabilistic bound
Retrained with new data, hyperparameters, or architectures
Table 1: Model-change assumptions and robustness formulations of the methods evaluated in this work, and the changed models used in their original evaluations.
Family
Construction
New initialization
Same data, architecture, and training configuration but different seed
Bootstrap
Bootstrap resampling of the training set
Data deletion
Random deletion of 1%, 5%, or 10% of the training observations
Data addition
Addition of 25%, 50%, 75%, or 100% of the reserved update pool
Label update
Random label flips for 1%, 5%, or 10% of the training observations
Training configuration
Prespecified combinations of optimizer, learning rate, and weight decay
Table 2: Construction of the held-out model variants. Each row contains 25 variants per base classifier. All families except new initialization reuse the base model’s initialization seed, so they isolate the data or training change from initialization noise.
Figure 1: Empirical robustness across held-out model changes. Panels correspond to datasets, rows to CFE methods, and columns to the eight change families. Each cell aggregates 25 variants for each of five independently trained base models and measures changed-model validity among base-valid CFEs.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
n
p
Split sizes
Factuals
Test balanced accuracy
Breast Cancer
569
30
284/57/57/171
312 (60–64)
0.967/0.973/0.988
Diabetes
768
8
384/76/77/231
281 (50–61)
0.712/0.721/0.736
Wine Quality
6,497
11
3248/649/650/1950
1,250 (250)
0.735/0.741/0.749
HELOC
8,291
20
4145/829/829/2488
1,250 (250)
0.727/0.730/0.735
Appendix
Table 3: Dataset and base-model summary. n and p denote the numbers of observations and input features. Split sizes are reported as training/update/validation/test. “Factuals” gives the total number of adverse test predictions evaluated across the five base models, followed by the per-model range in parentheses. Test balanced accuracy is reported as minimum/mean/maximum across the five models.
Change family
Hard disagreement (%)
Probability MAE
Bootstrap
4.92 ± 1.25
0.039 ± 0.010
Architecture
4.69 ± 1.54
0.038 ± 0.017
New initialization
4.57 ± 0.93
0.036 ± 0.005
Label update
3.69 ± 1.99
0.045 ± 0.025
Training configuration
3.10 ± 1.99
0.028 ± 0.019
Data addition
2.18 ± 0.96
0.017 ± 0.007
Appendix
Table 4: Output difference from the base classifier: mean ± standard deviation over the 125 variants per dataset, macro-averaged across datasets.
Method
Mean ± SD
Median
90th percentile
Timeouts
AP Δ S
3.909 ± 4.643
2.612
7.542
43
BetaRCE
0.0198 ± 0.0134
0.0159
0.0359
0
KD-tree
0.0037 ± 0.0022
0.0042
0.0050
0
RBR
1.754 ± 0.909
1.860
2.900
0
RNCE
0.262 ± 4.200
0.0007
0.0008
0
ROAR-LIME
0.086 ± 0.060
0.074
0.151
0
Appendix
Table 5: Per-factual CFE generation time in seconds, reported as mean ± standard deviation across requests. AP Δ S timeouts are unsuccessful requests that reach its 30-second limit.
Method
Coverage ↑
Base valid ↑
Emp. robust. ↑
E2E robust. ↑
Distance ↓
Breast Cancer
AP Δ S
90.4 ± 19.8
100.0 ± 0.0
58.5 ± 3.6
53.1 ± 12.8
0.689 ± 0.201
BetaRCE
100.0 ± 0.0
100.0 ± 0.0
97.0 ± 0.8
97.0 ± 0.8
1.860 ± 0.052
KD-tree
100.0 ± 0.0
100.0 ± 0.0
94.8 ± 2.4
94.8 ± 2.4
1.554 ± 0.006
RBR
100.0 ± 0.0
100.0 ± 0.0
77.1 ± 2.4
77.1 ± 2.4
1.097 ± 0.022
RNCE
100.0 ± 0.0
100.0 ± 0.0
98.5 ± 0.8
98.5 ± 0.8
1.571 ± 0.010
Appendix
Table 6: Full five-seed results, reported as mean ± standard deviation over the five base models.
Wrocław University of Science and Technology, Wrocław 50-347, Poland · Tooploox, Wrocław 53-601, Poland · Poznań University of Technology, Poznań 61-131, Poland