In the field of Explainable AI (XAI), counterfactual (CF) explanations interpret a model's decision by suggesting the changes to the input that would lead to a more favourable outcome. To be useful in practice, such an explanation should change few features and change them as little as possible, properties known as sparsity and proximity. We observe that existing methods remain limited in this respect, especially for numerical features, whether they are model-agnostic and amortised, or gradient-based with full access to the model. In this paper, we propose FlowCF, a model-agnostic generative method that frames CF generation as sparse transport from the factual to the target class. We solve this transport with flow matching, which we extend to mixed feature types with a novel mixed flow operator, and exploit the resulting geometry to optimise for sparsity through a gating network that minimises the number of features the transport changes. Extensive experiments on six benchmark datasets demonstrate that FlowCF produces the best numerical sparsity and proximity, changing 29% of the numerical features where the best baseline changes 89%, at 70% smaller displacement, while remaining comparable on the other desiderata.
Figures & tables
model- agnostic
amortised
exact numerical sparsity
no black-box queries at inference
Wachter
✗
✗
✗
✗
DiCE
✗
✗
✗
✗
TABCF
✗
✗
✗
✗
REVISE
✗
✗
✗
✗
CCHVAE
✓
✗
✗
✗
DiCoFlex
✓
✓
✗
✗
Table 1: Desired properties of existing counterfactual methods compared to FlowCF.
Figure 1: Counterfactuals generated by our method FlowCF vs DiCoFlex, the closest baseline method, on the two-moons dataset. FlowCF achieves exact numerical sparsity, changing a single feature when possible, while all counterfactuals sampled by DiCoFlex change both features.
Figure 2: FlowCF’s three training steps on samples from the 2D two-moons dataset. Green marks a learned displacement that changed one feature, red both. Step 1 produces valid CFs but changes both features. Step 2 achieves better sparsity making three of the four CFs axis-aligned. Step 3 further improves proximity while preserving validity and sparsity.
Dataset
Training size
#Features (Num/Cat)
Classification task
Adult
32,561
4 / 8
Income above or below $50K
Bank marketing
20,002
7 / 9
Subscription to a term deposit
Credit default
27,000
14 / 9
Default on the next payment
Lending Club
50,000
8 / 4
Loan fully paid or charged off
MAGIC
9,510
10 / 0
Gamma versus hadron air shower
HTRU2
8,949
8 / 0
Pulsar versus spurious candidate
Table 2: Dataset characteristics.
Validity (%) ↑
Sparsity (%) ↓
Proximity ↓
Top 2 ↑
Cat
Num
Num
Wachter
0.863
-
1.000
1.113
0
DiCE
1.000
0.145
0.893
0.957
15
REVISE
0.717
0.172
0.990
0.819
6
CCHVAE
0.900
0.196
1.000
0.763
6
TABCF
0.926
0.458
0.898
1.164
6
Table 3: Results averaged over all datasets. Top 2 counts best or runner-up finishes in Table 4 .
Validity (%) ↑
Sparsity (%) ↓
Proximity ↓
Validity (%) ↑
Sparsity (%) ↓
Proximity ↓
Cat
Num
Num
Cat
Num
Num
Adult
Bank marketing
Wachter
0.948
-
1.000
0.778
0.803
-
1.000
1.953
DiCE
1.000
0.011
0.771
2.643
1.000
0.164
0.822
1.347
REVISE
0.879
0.196
0.993
0.965
0.418
0.177
0.945
1.032
CCHVAE
0.784
0.213
1.000
0.685
0.934
0.216
1.000
0.591
Table 4: Detailed results on each dataset. Bold is best, underlined second best. Categorical sparsity does not apply to the continuous-only datasets MAGIC and HTRU2.
Figure 3: Counterfactuals generated by all methods for two CelebA-HQ instances, targeting the male and young attributes. The factual image is marked in green, and each image is annotated with the black-box’s probability P of the target class and the number of changed axes ∣Δaxes∣ . The bar plot gives the per-axis displacement in standard deviations, with our method FlowCF hatched.
Counterfactual (CF) explanations identify changes that alter an input's classification. While existing methods produce realistic and low-cost CFs, they often fail to ensure feasibility, by suggesting non-constructive modifications or incompatible with future changes (e.g., changing an individual's race to secure a job offer). We introduce a refinement of CF explanations that explicitly enforces feasibility. Our approach is the first to efficiently generate CFs that are realistic, low-cost and feasible. We accommodate both hard feasible constraints, specified by domain knowledge users, and soft feasible constraints, inferred automatically via causal inference from the dataset. Our method, Feasible Counterfactual Explanations (FCx), is based on a modified Variational Autoencoder (VAE) optimized with a multi-factor loss function. We measure the cost of a change based on the absolute change in values (proximity) as well as the number of features changed (sparsity) while realism is measured based on the LOF for density estimation, guaranteeing that CFs reside in densely populated regions. Extensive experiments on four public datasets show that our approach matches state-of-the-art performance across multiple metrics while guaranteeing feasibility.
Kleopatra Markou, Vana Kalogeraki, Dimitrios Gunopulos
National and Kapodistrian University of Athens, Department of Informatics and Telecommunications, Greece · Athens University of Economics and Business, Department of Informatics, Greece
Counterfactual explanations (CEs) are essential for actionable recourse, yet their reliability is often compromised in low-density regions, where classifiers exhibit high variance. Unlike existing methods that rely on expensive ensemble intersections to define stability, we propose \textit{DensityFlow}, a generative framework that constructs robust CEs by adhering to the high-confidence data manifold. Specifically, we model the counterfactual generation as continuous-time dynamics parameterized by Neural ODE, guided by a differentiable density score to actively avoid uncertain, low-density areas. This density score is learned via Noise Contrastive Estimation, effectively leveraging a (K+1)-way discriminator to estimate density ratios. For black-box settings, we introduce a local proxy distillation mechanism that aligns a lightweight surrogate with the target model strictly within the trajectory of CE generation, enabling efficient gradient-based optimization with minimal queries. Experiments demonstrate that \textit{DensityFlow} achieves superior validity under model multiplicity while significantly reducing query costs compared to ensemble-based baselines. Our implementation is available at https://github.com/G-AILab/DensityFlow.
Jun Tan, Qing Guo, Zicheng Xu +3
School of Computer Science and Engineering, Central South University, Changsha, China.
This paper proposes ConceptCF, a method for counterfactual generation that operates on human-interpretable concepts. In high-stakes domains such as healthcare and predictive maintenance, artificial intelligence models can increase efficiency and safety. Explainability is key to ensure these models rely on causal relationships rather than spurious correlations. Counterfactual explanations identify minimal modifications that would change a model's predictions. Existing methods for time series operate on individual points or subsequences without ensuring interpretability of the mutations. ConceptCF instead modifies meaningful concepts. As a result we can provide explanations in terms of these concepts, for example ``the model's prediction would be Sit' instead of Walk' if you increase the scale of the movement''. In this paper, the concepts are constructed through time series decomposition, resulting in concepts such as scale, and frequency bands. Counterfactuals are generated using a genetic algorithm that optimizes the concept mutations. Evaluation against five state-of-the-art approaches demonstrates that ConceptCF consistently achieves top-tier performance across validity, confidence, proximity, sparsity and plausibility metrics.
Annemarie Jutte, Faizan Ahmed, Jeroen Linssen +1
Saxion University of Applied Sciences, Enschede, The Netherlands · University of Twente, Enschede, The Netherlands