cs.SESep 20, 2025

CAFÉ: Causal Black-Box Testing of Machine Unlearning

Authors: Anna Mazhar, Sainyam Galhotra

Organizations: Cornell University, USA

Abstract

Machine learning models are increasingly deployed as software components that must evolve as requirements change. When specific training records or features must no longer influence a deployed model, machine unlearning aims to remove that influence without retraining from scratch. Because unlearning is often approximate, its effectiveness must be tested. Such tests must often treat the model as a black box, without access to its parameters, training history, or unlearning procedure. Features pose a further challenge: even after a feature is removed from a model's inputs, its influence can persist through downstream features. Many existing checks examine only the feature's direct use and can therefore certify a model that still depends on it. We frame unlearning testing as specification-based testing and present CAFÉ, which, using only a deployed model's predictions, intervenes on the feature, propagates the change to its downstream features, and checks whether the predictions still respond. CAFÉ measures a target's residual influence through both its direct and indirect causal paths, and its fine-grained diagnostics show which channels and subgroups still carry it. On two causal-network benchmarks with four unlearning methods, CAFÉ ranks residual influence with 0.92--0.93 pairwise accuracy, against at most 0.71 for existing checks, which fail in both directions: they certify models whose influence persists through downstream features and flag correctly unlearned ones. On real census data, CAFÉ likewise exposes influence that survives retraining yet goes unnoticed by direct-input checks.

Figures & tables

Explore similar work

CardsList
  1. Towards Reliable Testing of Machine Unlearning

    Apr 16, 2026Anna Mazhar, Sainyam GalhotraData LeakageAutomated Software Testing

  2. RULER: Representation-Level Verification of Machine Unlearning

    May 26, 2026Georgina Cosma, Axel FinkeMachine UnlearningMachine Unlearning Evaluation

  3. Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification

    Jul 21, 2026Sen Yang, Yuen-Hei YeungMachine UnlearningMachine Unlearning Evaluation