Organizations: BIFOLD, Germany · Technical University of Berlin, Germany · RIKEN AIP, Japan · Shanghai Jiao Tong University, China · The University of Sydney, Australia · University of Trento, Italy · Max Planck Institute for Intelligent Systems, Germany
Recent advances in representation learning have highlighted the utility of constant-curvature models, such as hyperbolic and spherical spaces, for modeling complex data. Mixed-curvature models further enhance this by integrating multiple constant-curvature components. However, these models typically learn each component space independently because spaces with different curvatures are inherently heterogeneous and lack a unified metric. Consequently, they lack explicit mechanisms to enforce geometric consistency across various spaces. Moreover, the problem of comparing probability distributions across mixed-curvature spaces remains unexplored. To compare distributions on heterogeneous spaces, Gromov-Wasserstein (GW) distances provide a principled framework by aligning their intra-space geometries. Building on this, we propose constant-curvature sliced Gromov-Wasserstein (CCSGW), a novel divergence for aligning distributions supported on heterogeneous constant-curvature spaces. We first introduce the missing geodesic-based one-dimensional projections for spherical spaces, and then extend sliced GW to constant-curvature spaces, enabling efficient and principled comparison across manifolds with different curvatures. This formulation preserves intrinsic geometric relationships while avoiding the high computational cost. We provide theoretical analysis showing that CCSGW controls intrinsic geometric discrepancy across heterogeneous spaces, promoting distribution-level geometric consistency. By integrating CCSGW into existing mixed-curvature learning tasks, including graph anomaly detection, graph node classification, and multimodal learning, we observe consistent performance gains across diverse settings.
Figures & tables
Figure 1: CCSGW Overview. Distributions supported on spherical, Euclidean, and hyperbolic spaces are projected onto one-dimensional coordinates using geometry-aware projections. CCSGW then compares the projected distributions across heterogeneous constant-curvature spaces, encouraging consistency of their intrinsic geometric structures despite differences in curvature.
Figure 2: Simulation of cross-curvature alignment. A Euclidean point cloud initialized as a line is optimized to align with a fixed ring-shaped distribution in hyperbolic space using CCSGW. As optimization proceeds, the Euclidean points gradually form a circular structure while the CCSGW discrepancy decreases, illustrating improved geometric consistency.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Hyperparameter sensitivity of CCSGW. Each panel varies one hyperparameter, namely the alignment weight λ , the number of geodesic directions L or the number of sampled nodes n (“all” uses every node), while the other two are fixed at the reference configuration (filled markers). The vertical axis shows the gain Δ of the backbone trained with CCSGW over the same backbone trained without it (gray line at zero). Points and error bars are the mean and 95% confidence interval over ten splits.
Figure 4: Time of one alignment-loss evaluation (forward and backward) as a function of the number of sampled nodes n , on a single A100 GPU (median of seven runs, log–log axes). GW runs out of memory beyond n=4096 .
The Gromov-Wasserstein (GW) problem provides a framework for aligning heterogeneous datasets by matching their intrinsic geometry, but its statistical and computational scaling remains an issue for high-dimensional problems. Slicing techniques offer an appealing route to scalability, but, unlike Wasserstein distances, GW problems do not generally admit closed-form solutions in one-dimension. We resolve this problem for the GW problem with inner product cost (IGW), propose a sliced IGW distance that enjoys a natural rotational invariance property, and comprehensively study its structural and computational properties. Numerical experiments validating our theory are presented, followed by applications to heterogeneous clustering of text data and language model representation comparison.
Principal Component Analysis (PCA) is a fundamental tool for representation learning, but its global linear formulation fails to capture the structure of data supported on curved manifolds. In contrast, manifold learning methods model nonlinearity but often sacrifice the spectral structure and stability of PCA. We propose \emph{Geodesic Tangent Space Aggregation PCA (GTSA-PCA)}, a geometric extension of PCA that integrates curvature awareness and geodesic consistency within a unified spectral framework. Our approach replaces the global covariance operator with curvature-weighted local covariance operators defined over a k-nearest neighbor graph, yielding local tangent subspaces that adapt to the manifold while suppressing high-curvature distortions. We then introduce a geodesic alignment operator that combines intrinsic graph distances with subspace affinities to globally synchronize these local representations. The resulting operator admits a spectral decomposition whose leading components define a geometry-aware embedding. We further incorporate semi-supervised information to guide the alignment, improving discriminative structure with minimal supervision. Experiments on real datasets show consistent improvements over PCA, Kernel PCA, Supervised PCA and strong graph-based baselines such as UMAP, particularly in small sample size and high-curvature regimes. Our results position GTSA-PCA as a principled bridge between statistical and geometric approaches to dimensionality reduction.
Alexandre L. M. Levada
Computing Department Federal University of São Carlos 13565-905, São Carlos-SP, Brazil
We introduce Convex Distance Operator Transport (CDOT), the first convex optimal transport framework that aligns distributions across heterogeneous domains by jointly preserving feature correspondence and intrinsic geometric structure. Specifically, CDOT employs an operator-based regularization that aligns aggregated distance structures by introducing distance and conditional expectation operators. Consequently, the proposed regularization improves the robustness to local geometric variations. We further prove that the resulting CDOT discrepancy is a valid pseudometric on the space of attributed compact metric-measure spaces. In addition, we characterize the relationship between CDOT and Gromov--Wasserstein (GW) through a new notion of dispersion gap, formally elucidating the geometric source of non-convexity in GW compared to the convexity of CDOT. In the finite-sample regime, we derive a non-asymptotic risk bound decomposed into optimization and statistical errors, establishing risk consistency under a globally convergent Frank--Wolfe algorithm. Experiments on synthetic point clouds, brain connectomes, and graph classification benchmarks demonstrate better performance over existing methods, with stable and reliable behavior in practice.
Junhyoung Chung, Euijong Song, Won Hwa Kim +1
KRAFTON · Department of Statistics, Seoul National University, Seoul, South Korea · Graduate School of AI, POSTECH, Pohang, South Korea +3