cs.CVSep 27, 2026

M3-Score: Fidelity, Memorization and Coverage as Separate Axes for Evaluating Generative Radiology Image Models

Authors: Sathiyamohan Nishankar, Pubudu Sanjeewani, Asanka Perera

Organizations: Faculty of Engineering, University of Peradeniya, Sri Lanka · School of Computing Technologies, RMIT University, Melbourne, Australia · School of Engineering & Digital Technologies, University of Southern Queensland, Brisbane, Australia · School of Engineering & Technologies, UNSW, Canberra, Australia

Abstract

Quantitative evaluation of generative models for radiology remains challenging. Clinically relevant structures are often small and infrequent, feature spaces learned from natural images may represent them poorly, and a single summary score cannot distinguish limited fidelity from limited diversity. This study proposes the Medical Multi-axis Maximum Mean Discrepancy score (M3-Score), an evaluation framework based on RadioDINO-s16, a frozen vision transformer pretrained on radiology images. M3-Score reports three complementary axes computed at pre-specified encoder depths: \emph{fidelity}, measured by an unbiased multi-bandwidth radial basis function (RBF) MMD2^2 at the final block; \emph{memorization}, measured by nearest-neighbor distances at 75% depth; and \emph{coverage}, defined as the fraction of real images with a generated neighbor within their kk-nearest-neighbor radius at 33% depth. Reference sets are sampled across subjects to limit the influence of correlated slices. On BraTS brain MRI, the fidelity axis ordered five comparison sets of increasing severity (Spearman ρ=1.00ρ= 1.00), and a subject-disjoint real set yielded MMD2=0\mathrm{MMD}^2 = 0 (permutation p=1p = 1). An unconditional denoising diffusion probabilistic model achieved MMD2=0.073\mathrm{MMD}^2 = 0.073 (95% confidence interval [0.071,0.080][0.071, 0.080]) but covered only 38% of the real distribution. Under progressive mode dropping, kk-NN manifold recall increased at all twelve encoder blocks, whereas the proposed coverage estimator decreased monotonically (ρ=−1.00ρ= -1.00). RadioDINO-s16 features separated real brain MRI from generated samples with a ROC-AUC of 0.819, compared with 0.555 for InceptionV3 and 0.582 for CLIP. Across a twentyfold range of sample sizes, the mean M3 value varied by a factor of 1.05, compared with 2.52 for the Fréchet Inception Distance.

Figures & tables

Explore similar work

CardsList
  1. MedForj: An open, large-scale foundational generative prior for high-resolution 3D brain MRI

    Oct 29, 2025Samuel W. Remedios, Aaron Carass, Jerry L. Prince +1Medical Image GenerationDiffusion Priors

  2. Tokenizer-Generator Coupling in Medical Image Generation

    Aug 7, 2026Liam ChalcroftMedical Image GenerationVisual Tokenizers

  3. Scaling Generative Foundation Models for Chest Radiography with Rectified Flow Transformers

    Jun 17, 2026Fabio De Sousa Ribeiro, Emma A. M. Stanley, Charles Jones +7Digitally Reconstructed RadiographsMedical Image Generation