Machine learning models are increasingly deployed as software components that must evolve as requirements change. When specific training records or features must no longer influence a deployed model, machine unlearning aims to remove that influence without retraining from scratch. Because unlearning is often approximate, its effectiveness must be tested. Such tests must often treat the model as a black box, without access to its parameters, training history, or unlearning procedure. Features pose a further challenge: even after a feature is removed from a model's inputs, its influence can persist through downstream features. Many existing checks examine only the feature's direct use and can therefore certify a model that still depends on it. We frame unlearning testing as specification-based testing and present CAFÉ, which, using only a deployed model's predictions, intervenes on the feature, propagates the change to its downstream features, and checks whether the predictions still respond. CAFÉ measures a target's residual influence through both its direct and indirect causal paths, and its fine-grained diagnostics show which channels and subgroups still carry it. On two causal-network benchmarks with four unlearning methods, CAFÉ ranks residual influence with 0.92--0.93 pairwise accuracy, against at most 0.71 for existing checks, which fail in both directions: they certify models whose influence persists through downstream features and flag correctly unlearned ones. On real census data, CAFÉ likewise exposes influence that survives retraining yet goes unnoticed by direct-input checks.
Figures & tables
Figure 1 . The causal specification describes the real-world outcome; the probes test the deployed model’s prediction of that outcome. Changing location while freezing all other inputs produces no model response, but propagating the intervention through commute distance changes the prediction. Income remains fixed because its association with location is due only to a common cause.
Retrained without target column ↑
All candidates ↑
False alarms on clean controls ↓
Direct-input probes ∗
0.00 / 0.00
0.75 / 0.75
0.00 / 0.00
Backdoor adjustment
0.47 / 0.00
0.50 / 0.02
0.00 / 0.07
Statistical parity difference
0.27 / 0.00
0.18 / 0.00
0.40 / 0.00
Attribute inference attack
0.13 / 0.07
0.27 / 0.12
0.33 / 0.13
CAFÉ
1.00 / 1.00
1.00 / 1.00
0.00 / 0.00
Table 1 . Strict detection ( δ=0 ) rates, as Insurance / ALARM. Per network: 60 candidates, 15 of them retrained without the target’s column, and 15 clean controls. CAFÉ and the direct-input probes fire when a bootstrap lower bound on mean absolute influence exceeds zero; the statistical baselines, when their statistic exceeds 95% of a permutation null. Detection asks only whether influence is nonzero; ranking (Figure 2 ) is the harder test.
Figure 2 . Pairwise ranking accuracy of every feature’s estimated residual influence against its reference effect (0.5 is chance), after each unlearning method, on Insurance and ALARM. One mark per unlearning method, with the bar showing ± 1 SD over 5 seeds and the grey spine the spread across methods. The attribute inference attack has no network-wide ranking: rare feature values leave its attacker too few examples to train on. CAFÉ ranks most accurately after every unlearning method, on both networks.
Figure 3 . Left: channel ranking accuracy on the benchmark networks, whose channels share mediators (0.5 is chance); feature-scoring baselines stay at chance. Right: each mark is one model; it shows how much influence is carried by the channel a method picks as strongest, relative to the true strongest channel (1 means the method picked it). On ALARM, where channels nearly tie, CAFÉ ’s picks stay close to 1.
Propagation
Insurance
ALARM
Fresh noise
0.914 / 0.0113
0.913 / 0.0062
Conditional mean
0.861 / 0.0233
0.875 / 0.0153
Own noise ( CAFÉ )
0.919 / 0.0105
0.932 / 0.0058
Table 2 . Propagation variants with the graph and fitted mechanisms held fixed: pairwise ranking accuracy / effect MAE. All three detect every candidate with no false alarm on the clean controls.
Network
5%
10%
20%
Insurance
0.928 / 0.0162
0.914 / 0.0248
0.890 / 0.0298
ALARM
0.944 / 0.0125
0.918 / 0.0172
0.910 / 0.0147
Table 3 . Rank accuracy / effect MAE under random edge corruption (feature shuffling). Matching correct-graph references are 0.936 / 0.0151 on Insurance and 0.951 / 0.0079 on ALARM.
Figure 4 . Pairwise ranking accuracy over all features against the queries each audit actually used. CAFÉ ’s curve is the budget sweep, one point per cap, plotted at its measured cost, which falls below the cap; each comparator is one mark at the largest predeclared test set it can afford, with ± 1 SD over 5 seeds. Statistical parity difference is the cheapest method in the comparison at 480 queries, but its accuracy is fixed: buying more budget cannot improve a group rate gap that already uses every test instance. The attribute inference attack is omitted: it has no network-wide ranking (Section 5.2 ). CAFÉ ranks most accurately at every budget.
Machine learning components are now central to AI-infused software systems, from recommendations and code assistants to clinical decision support. As regulations and governance frameworks increasingly require deleting sensitive data from deployed models, machine unlearning is emerging as a practical alternative to full retraining. However, unlearning introduces a software quality-assurance challenge: under realistic deployment constraints and imperfect oracles, how can we test that a model no longer relies on targeted information? This paper frames unlearning testing as a first-class software engineering problem. We argue that practical unlearning tests must provide (i) thorough coverage over proxy and mediated influence pathways, (ii) debuggable diagnostics that localize where leakage persists, (iii) cost-effective regression-style execution under query budgets, and (iv) black-box applicability for API-deployed models. We outline a causal, pathway-centric perspective, causal fuzzing, that generates budgeted interventions to estimate residual direct and indirect effects and produce actionable "leakage reports". Proof-of-concept results illustrate that standard attribution checks can miss residual influence due to proxy pathways, cancellation effects, and subgroup masking, motivating causal testing as a promising direction for unlearning testing.
Machine unlearning aims to remove the influence of specific training records from a deployed model without retraining from scratch. Current protocols verify this at the output level through membership inference, retain accuracy, and forget-set accuracy, but a model can satisfy all three whilst still encoding forgotten records in its intermediate representations. We introduce RULER, a set of representation-level verification metrics. The oracle-comparative metric M2 measures whether forget-set records occupy the same representational position as in a model retrained without them. The oracle-free metric M4 detects residuals from the unlearned model's internal similarity structure alone, without retraining. Four approximate unlearning methods all pass output-level evaluation, yet under a linear mixed-effects model M2 detects significant residuals in 10 of 12 conditions (p<0.05), with effect sizes growing as the forget fraction increases. A fifth method, Bad Teacher, shows the same residuals despite a different forgetting mechanism. M4 acts as a pre-unlearning diagnostic across tabular, image, clinical text, and face-identity settings: it detects identity-level memorisation in face recognition models where no tested method fully erases the signal.
Georgina Cosma, Axel Finke
Department of Computer Science, Loughborough University, UK · School of Mathematics, Statistics and Physics, Newcastle University, UK
Machine unlearning is commonly evaluated by matching a retrained oracle on trained probes. In a controlled nonce-fact testbed with a matched retraining reference, we find this criterion can favor methods that retain held-out knowledge: candidates it rates adequate score held-out forget facts −2.82 nats below the never-learned level (cluster CI [−3.16,−2.48]). We recast unlearning as restoration to the matched reference and audit oracle-free screens and certificate-style criteria across 45 model-seed cells spanning five open architecture families. The reference itself falsifies an absolute retain/round-trip certificate: the injected model, which retains the retain set by construction, fails the fixed retain threshold in 41/45 cells and its own round trip in 31/45, and the reference fully certifies in only 1/45. A base-anchored held-out screen remains strong as a selective necessary test: on a sealed challenge suite it rejects the injected model in 45/45 cells, accepts the reference in 44/45, and partially detects entity-routing suppression (35/45); it is a necessary test with measured sensitivity, not a sufficiency certificate. A damage-relative recalibration anchored to the reference's own operating point certifies a small subset in 15/45 cells; where it does not abstain, its picks lie within retraining noise (0.80 nats) on the axes it optimizes, while the common trained-probe criterion sits 5.17 nats away (a supporting comparison, not a head-to-head benchmark). A fixed-magnitude logit-suppression attack defeats the full forward battery in 12/45 cells, so forward-only certification is not sound; our method is an empirical selective test for methods-as-produced. An identifiability theorem delimits which facts admit an oracle-free forget threshold at all, with TOFU as the predicted boundary case.
Sen Yang, Yuen-Hei Yeung
Stern School of Business New York University · Courant Institute of Mathematical Sciences New York University