False negatives remain a critical limitation of computer-aided diagnosis (CAD) systems for breast cancer screening due to delayed detection and treatment. To address this issue, we propose a counterfactual data augmentation strategy that generates healthy mammograms by "erasing" lesions from anomalous images, thereby enriching the training distribution. We train a Denoising Diffusion Probabilistic Model on BI-RADS 1 (healthy) mammograms and use a RePaint-based sampling strategy to inpaint realistic normal tissue within annotated lesion bounding boxes. The resulting healthy counterfactuals replace annotated lesion regions with realistic healthy tissue while preserving patient-specific anatomical structure, as supported by similarity metrics between real and generated images. Image realism was further assessed by radiologists and found to be consistent with the original dataset quality. We evaluate counterfactual augmentation across four representative classifier architectures: a convolutional neural network (ConvNeXt), a vision transformer (ViT), a vision-language model pre-trained on mammogram-report pairs (Mammo-CLIP) and a multi-scale attention-based multiple-instance learning framework (FPN-MIL). Experiments conducted on the VinDr-Mammo dataset show improvements in sensitivity across all architectures, particularly at 80% fixed specificity, contributing towards more reliable CAD systems for breast cancer. Code is available at: https://github.com/ines03garcia/diffusion-based-counterfactual-generation.
Figures & tables
Figure 1: Overview of the proposed pipeline. First, a DDPM is trained on healthy mammograms to learn the distribution of healthy breast tissue. Second, RePaint is used to generate healthy counterfactuals via mask-guided inpainting of lesion regions, using binary masks derived from bounding-box annotations. Finally, the generated counterfactuals are incorporated into the training set for classification.
Compared Distributions
Dataset
Scale
FID ↓
FRDv0 ↓
FRDv1 ↓
Healthy CF. vs. Real Healthy (ours)
VinDr-Mammo
image
10.1
12.3
7.5
Synthetic vs. Real [ 5 ]
VinDr-Mammo
image
67.5
–
12.4
Healthy CF. vs. Real Healthy (ours)
VinDr-Mammo
patch
37.7
12.3
10.9
Synthetic Masses vs. Real Masses [ 20 ]
CBIS-DDSM
patch
58.0
18.1
–
Table 1: Quantitative comparison of image realism, measured by FID, and radiomic consistency, measured by FRD, between generated and real mammograms. Lower values indicate greater similarity to the corresponding real-image distribution. Results from the proposed method are shown in bold. Prior results are included for reference, although they are not directly comparable due to differences in protocol.
Figure 2: Distribution of radiologist ratings for real healthy mammograms and generated healthy counterfactuals. Ratings were assigned on a five-point Likert scale. For the full-image assessment, higher scores indicated that an image was more likely to resemble a real mammogram; for the bounding-box assessment, higher scores indicated that the highlighted tissue was more likely to resemble healthy breast tissue.
Figure 3: Three lesion-free counterfactual mammograms rated 5/5 for realism by all three reviewing radiologists. For each example, the original image is shown on the left and the corresponding counterfactual on the right. Zoomed views of the regions of interest are shown in the bottom row.
Figure 4: Radiologist evaluation of counterfactual realism for a balanced subset of counterfactuals. Left: Boxplot showing the distribution of mean realism ratings across breast density categories. Right: Scatter plot with linear regression illustrating the correlation between mean realism ratings and annotation area.
ConvNeXt [ 14 ]
ViT [ 3 ]
Mammo-CLIP [ 19 ]
FPN-MIL [ 17 ]
Metric
Baseline
CF
Baseline
CF
Baseline
CF
Baseline
CF
Balanced Acc.
73.1
73.2
72.3
72.7
77.1
77.3
76.7
76.8
Recall
63.5
68.1
63.1
63.7
69.0
70.7
73.1
75.0
Specificity
82.6
78.2
81.6
81.7
85.2
83.8
80.2
78.6
F1-score
63.9
64.1
62.9
63.3
69.3
69.4
68.6
68.7
ROC-AUC
80.3
80.2
77.8
78.4
84.6
84.8
84.1
84.0
Table 2: Classification performance of ConvNeXt, ViT, Mammo-CLIP and FPN-MIL under baseline training and counterfactual augmentation (CF) on the original test set. Improved results are highlighted in bold. Recall values with fixed specificity (80%) are also reported.
Counterfactual image generation answers questions about how a subject would have looked under retrospective, hypothetical scenarios. Recent methods have improved perceptual quality, identity preservation and faithfulness to an underlying causal model, but their adoption in healthcare is limited by scarce annotated data, distribution shift between datasets, and mismatches between pretrained generative models and those required for counterfactual inference. We propose specialisation, a data and parameter-efficient framework for adapting pretrained, non-causal generative models into causal mechanisms under distribution shift. Based on this framework, we train a radiology counterfactual image generation model, called RadCF, using latent flow matching. We validate our approach on three chest X-ray datasets spanning different dataset shifts, data volumes, and counterfactual questions, associated with challenging, highly-localised interventions. Our results show that RadCF and specialisation improve counterfactual soundness over existing methods while being data and parameter efficient, and that the resulting counterfactuals can detect and mitigate shortcut learning in a downstream medical classifier. Code is available at https://github.com/GSK-AI/RadCF/.
Ascription of an image gives insights into the objects that influence the classification of the whole image or its pixels towards a specific category. These insights help radiologists to visualize deformities in medical imaging. Most of the existing visualization techniques are based on discriminative models and highlight regions of the input image participating in the decision-making of a classifier. However, these approaches do not take all noticeable objects into account as their objective is to classify the input by using a minimal set of discriminative features. To overcome the issue, a counterfactual explanation (CX) based class-oriented feature attribution method is proposed. A counterfactual explanation (CX) explicates a causal reasoning process of the form: "if X had not happened, then Y would not have happened". The method is built on generative adversarial networks (GANs) with a cyclical-consistent loss function. We evaluate our method on three datasets: synthetic, tuberculosis and BraTS. All experiments confirm the efficacy of the proposed method. This study also highlighted the limitations of existing counterfactual explanation techniques in producing plausible counterfactual instances (CIs). Accompanying CXs with believable CIs thus provides self-explanatory analogy-based explanations. To this end, a CI generation method is proposed. Also, a novel technique is used to evaluate the quality of CI. The baseline results are produced on the BraTS dataset.
Generative augmentation is often proposed as a remedy for small medical-image datasets, but synthetic images are only useful when they improve downstream task performance. "Augmentation" here means synthetic supplementation: GAN-generated samples added to the real training pool, not geometric or photometric transforms of existing images. Twelve class-plane StyleGAN2-ADA generators were trained on constrained BRISC 2025 partitions to test whether their output, with or without InceptionV3 feature-space filtering, improves held-out tumour classification across three classifier families: a random forest (RF) on InceptionV3 features, a compact two-headed convolutional neural network (CNN), and MobileViTV2, a mobile hybrid convolutional-transformer. Each was evaluated at 1:1 and 1:2 real-to-synthetic ratios. An independent GPT-5.5 blind test placed gated real-versus-synthetic discrimination at 57.73% (95% CI: 54.48--60.92%) on the model-legible subset -- modestly above chance. The RF classifier did not benefit from the synthetic MRIs. The CNN showed consistent mean gains that did not survive Holm correction. MobileViTV2 showed the clearest benefit: filtered 1:1 augmentation improved tumour classification accuracy by 1.02% absolute (95% CI: 0.54--1.54%; Holm-corrected p = 0.0104). A secondary efficiency analysis found that every augmented CNN condition selected its checkpoint 42--64% earlier than baseline, while compute-matched MobileViTV2 runs reached selection after 50--67% fewer real-data epochs. Overall, augmentation utility was found to be architecture- and ratio-dependent, not guaranteed by visual fidelity alone.