cs.LGSep 2, 2025

Differentiable Expectation-Maximisation and Applications to Gaussian Mixture Model Optimal Transport

Authors: Samuel Boïté, Eloi Tanguy, Julie Delon, Agnès Desolneux, Rémi Flamary

Organizations: Université Paris Cité, CNRS, MAP5, F-75006 Paris, France · Centre Borelli, CNRS and ENS Paris-Saclay, F-91190 Gif-sur-Yvette, France · CMAP, CNRS, Ecole Polytechnique, Institut Polytechnique de Paris

Abstract

The Expectation-Maximisation (EM) algorithm is a central tool in statistics and machine learning, widely used for latent-variable models such as Gaussian Mixture Models (GMMs). Despite its ubiquity, EM is typically treated as a non-differentiable black box, preventing its integration into modern learning pipelines where end-to-end gradient propagation is essential. In this work, we present and compare several differentiation strategies for EM, from full automatic differentiation to approximate methods, assessing their accuracy and computational efficiency. As a key application, we leverage this differentiable EM in the computation of the Mixture Wasserstein distance MW2\mathrm{MW}_2 between GMMs, allowing MW2\mathrm{MW}_2 to be used as a differentiable loss in imaging and machine learning tasks. To complement our practical use of MW2\mathrm{MW}_2, we contribute a novel stability result which provides theoretical justification for the use of MW2\mathrm{MW}_2 with EM, and also introduce a novel unbalanced variant of MW2\mathrm{MW}_2. Numerical experiments on barycentre computation, colour and style transfer, image generation, and texture synthesis illustrate the versatility of the proposed approach in different settings.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Averaged Mirror Descent and Dual Gradient Methods: Convergent Algorithms for Entropic Gromov-Wasserstein Problems

    Sep 25, 2026Joanna Marks, Gabriel Rioux, Riccardo PasseggeriGromov--WassersteinMirror Descent

  2. On the Wasserstein Gradient Flow Interpretation of Drifting Models

    May 6, 2026Arthur Gretton, Li Kevin Wenliang, Alexandre Galashov +3Wasserstein Gradient FlowsGenerative Models

  3. Distance-Matrix Wasserstein Statistics for Scalable Gromov--Wasserstein Learning

    May 14, 2026Ao Xu, Tieru WuGromov--WassersteinWasserstein Distance