astro-ph.COJun 9, 2026

Interpretable Neural Marked Statistics for Cosmological Inference

Authors: Federico SemenzatoBenjamin D. WandeltMichele LiguoriAlvise Raccanelli

Abstract

Recovering cosmological information beyond the power spectrum is a central goal for upcoming cosmological surveys, since late-time non-Gaussian signal in the matter density cannot be accessed through two-point statistics alone. Marked statistics fold part of this information back into the two-point level by reweighting the field with non-linear functions. We propose a neural marking scheme to generalize this process through a set of interpretable, physically motivated transformations that directly allow to interpret the gain in cosmological information at the morphological level. We employ a contrastive learning objective to align learnable marked summaries with the underlying cosmological parameters. At kmax=0.2hMpc1k_{\max}=0.2\,h\mathrm{Mpc}^{-1}, our neural mark tightens the marginalized constraint on σ8σ_8 by 2.9×2.9\times and on ΩmΩ_m by 1.8×1.8\times compared to classical marks, breaking the Ωmσ8Ω_m-σ_8 degeneracy at the Fisher information level. It further reduces the parameter MSE across our cosmological parameter prior by 1.45×1.45\times over the best classical mark. The learned latent geometry aligns with the ΩmΩ_m and σ8σ_8 directions in parameter space, indicating that the contrastive objective recovers the dominant axes of cosmological information. Our approach opens the door to more powerful, interpretable summary statistics for cosmological inference.

Explore similar work

Jun 8, 2026astro-ph.CO

Learning the Universe: Posterior Reliability of Neural Generative Models in High-Dimensional Field-Level Inference of Cosmic Initial Conditions

Accurate posterior estimation is central to scientific inference, as uncertainties determine what can be reliably learned from observational data. While Markov chain Monte Carlo methods provide asymptotic convergence guarantees, they are computationally demanding in high-dimensional settings. Neural network-based generative models for entire discretized 3D fields enable fast amortized inference but often lack convergence guarantees and principled accuracy assessment. Using Hamiltonian Monte Carlo to obtain reference posterior samples, we conduct a controlled field-level evaluation of an implicit generative model (Stochastic Interpolants) and an explicit likelihood-based model (GLOW normalizing flows). This comparison, unavailable in typical applications, enables the detection of posterior geometry failures that standard metrics cannot capture. As a case study, we consider the cosmological inverse problem of inferring cosmic initial conditions from present-day large-scale structure. To match the precision of modern cosmological data, this problem increasingly relies on complex, non-linear, and non-differentiable simulators, which are incompatible with gradient-based inference frameworks. Generative models offer a route to address these challenges, provided their inferred posteriors are reliable. In this work, we show that matching posterior means, marginal distributions, or achieving high cross-correlation does not imply correct uncertainty structure, as revealed by posterior variance fields and sample-based evaluations. Through this work, we aim to raise awareness of the challenges of uncertainty estimation in high-dimensional field-level settings, highlighting the importance of careful design and validation of neural generative approaches for scientific applications.
Ludvig Doeser, Jens Jasche
Sep 8, 2026astro-ph.CO

Inductive Biases in Field-Level Cosmological Inference from Galaxy Catalogs

We perform field-level likelihood-free inference of the matter density parameter ΩmΩ_m from simulated galaxy catalogs using machine learning models with differing inductive biases. Using hydrodynamic simulations from CAMELS, we examine how observable choice and architecture govern cosmological information extraction. We consider galaxy positions and line-of-sight peculiar velocities, separately and jointly, and compare permutation-invariant Deep Sets, implemented with either multilayer perceptrons (MLPs) or Kolmogorov-Arnold Networks (KANs), to graph neural networks (GNNs), which explicitly encode spatial relations. We test in-distribution and out-of-distribution (OOD) performance across simulations with different subgrid galaxy-formation prescriptions. Deep Sets infer ΩmΩ_m from velocities alone with mean relative errors of approximately 18%18\% in-distribution and 25%\sim25\% OOD, with KANs and MLPs achieving comparable performance. In contrast, the same set-based approach does not yield useful σ8σ_8 predictions in either in-distribution or cross-suite tests. Adding positions does not improve Deep Sets, while GNNs infer ΩmΩ_m with mean relative errors of about 10%10\% in-distribution and 1010--17%17\% OOD. These results indicate that peculiar velocities provide the dominant source of ΩmΩ_m information for set-based models in this setting, while spatial information is most effectively used by architectures that explicitly encode galaxy-galaxy relations. Because the velocity inputs are exact simulated peculiar velocities, applications to survey data will require validation under realistic velocity-measurement noise, selection effects, and survey geometry.
James O. Baldwin, Shy Genel, Francisco Villaescusa-Navarro
Sep 7, 2026astro-ph.IM

Neural Posterior Estimation for Tomographic Weak Lensing Mass Mapping

Weak gravitational lensing shear and convergence trace the distribution of baryonic and dark matter across space, making them a powerful probe of cosmic structure. Inferring shear and convergence from images is a challenging inverse problem. The prevailing approach to this task estimates shear from weighted averages of galaxy ellipticities, calibrates these estimates to account for systematic biases, and transforms them to reconstruct convergence, a multistage procedure that requires substantial computational resources and meticulous handling of statistical uncertainties. As an alternative, we propose a probabilistic approach to field-level weak lensing inference in which we train a deep neural network to directly map a multiband image to a variational distribution over the underlying tomographic shear and convergence fields. This neural posterior estimation (NPE) procedure implicitly marginalizes over nuisance variables in the cosmological forward model and does not require evaluating the likelihood function. It is also amortized, so it enables rapid posterior inference for astronomical surveys once the neural network is trained. When evaluated on synthetic images from the LSST-DESC DC2 Simulated Sky Survey, NPE produces well-calibrated variational distributions for shear and convergence that are consistent with the ground truth. We describe how maps sampled from these variational distributions could be used in a subsequent simulation-based inference procedure to approximate the posterior distribution over cosmological parameters.
Tim White, Shreyas Chandrashekaran, Camille Avestruz +2