The Degeneracy Distillery
Organizations: Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Wilberforce Rd, Cambridge CB3 0WA, United Kingdom · Imperial Centre for Inference and Cosmology (ICIC), Imperial College London, Prince Consort Road, London SW7 2AZ, United Kingdom · Astrophysics, University of Oxford, Oxford OX1 3RH, United Kingdom · CNRS & Sorbonne Université, Institut d’Astrophysique de Paris (IAP), UMR 7095, 98 bis bd Arago, F-75014 Paris, France · Department of Physics and Astronomy, University College London, London WC1E 6BT, United Kingdom · Department of Physics & King’s Institute for Artificial Intelligence, King’s College London, London WC2R 2LS, United Kingdom · Department of Physics and Astronomy, Johns Hopkins University, Baltimore, MD 21218, USA · Department of Applied Mathematics and Statistics, Johns Hopkins University, Baltimore, MD 21218, USA
Abstract
When two or more parameters or labels produce similar data, they are degenerate, or hard to distinguish. Degeneracies render both label prediction and inverse problems difficult, since both machine learning algorithms and probabilistic samplers rely on the distinguishability of data and its gradients with respect to parameters. However, identifying degeneracies in physical models or real-world datasets can be elucidating about the choice of model or the underlying process that produces the data. We present the degeneracy distillery, a method that (1) detects and (2) resolves degenerate parameter combinations (a) automatically and (b) symbolically, from parameter-data (or parameter-simulation) pairs alone, through estimation and flattening of the Fisher information matrix. By exploring the information geometry of the likelihood, we characterize degeneracies as an intrinsic property of the physical model, requiring no realised data observation. We demonstrate our approach on a range of synthetic and real-world problems, discovering symbolic coordinate transformations that identify the combinations of parameters of a model which yield independent effects on the data. The resulting coordinates flatten the Fisher information in expectation globally, in contrast to posterior-based methods that flatten only at a single point, and substantially reduce the simulation budget required for downstream neural posterior estimation. In test cases we require up to fewer simulations for posterior estimation at matched validation calibration whilst simultaneously gaining physical insight on the system.