We study the implicit bias of Riemannian gradient flow for hyperbolic multiclass classification with fixed class prototypes in hyperbolic space Hn. Our framework accommodates general permutation invariant relative margin (PERM) losses, a class that includes cross entropy and other standard multiclass losses. Our analysis is based on a decomposition: at large radius, the distance to each prototype splits into a radial term and a direction-dependent term described by the Busemann function. This yields two main results. First, we prove a radial dichotomy: the sign of a drift coefficient μ determines whether the radius is pushed toward the ideal boundary or back toward the interior; if the positive drift persists, then r(t)=21logt+O(1), while persistent negative drift returns the trajectory to the large-radius threshold in finite time. Second, we show that the boundary direction converges to a critical point of the Busemann risk on ∂Hn. These results provide a rigorous asymptotic perspective on two phenomena we refer to as boundary saturation and near-boundary clustering in hyperbolic representation learning.
Figures & tables
Figure 1 : Geometric illustration of hyperbolic representations and their boundary asymptotics.
Figure 2 : Radial dichotomy and boundary convergence.
Figure 3 : Asymptotic decision boundary on the ideal boundary.
Figure 4 : Numerical verification of radial dichotomy and angular convergence. Top: radial trajectories with positive drift align with the slope 21 reference in ri versus logt , while negative drift trajectories remain bounded. Bottom: angular trajectories converge to Busemann critical points on the ideal boundary. Ablations over prototype number, PERM template, and prototype geometry show the same qualitative behavior.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Drift coefficient μ(ξ^(t),y) along the angular trajectories. Blue: true-min attractors; orange: antipode attractors; dashed line: μ=0 .
Figure 6 : Poincaré-disk versions of the 4×5 summary in Fig. 4 . Top two rows: Theorem 6.2 radial trajectories ( ξ^i held fixed, so each ray is a straight line from the origin). Bottom two rows: Theorem 6.3 angular trajectories rendered with synthetic radial spread for visibility on S1 .
Hyperbolic geometry has recently emerged as a powerful representation space for multimodal learning, as it naturally captures hierarchical semantic structure across modalities. Despite this progress, how such representations behave under continual learning poses fundamentally different challenges that remain underexplored. This work provides a geometric perspective on this problem and establishes a theoretical foundation for representation preservation in hyperbolic space, showing that preventing forgetting requires cross-modal invariance under a shared hyperbolic isometry. We further show that forgetting in hyperbolic continual learning involves both semantic relation drift and hierarchy-related distortion, motivating preservation of both cross-modal relational structure and hierarchical geometry. Guided by these insights, a principled continual learning framework is derived that preserves essential geometric structure while allowing effective adaptation to new tasks. Experiments on continual multimodal benchmarks corroborate the effectiveness of the proposed approach.
Jiahong Liu, Ming Shen, Xiaohao Liu +4
Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong, China · School of Computing, National University of Singapore, Singapore · Department of Computer Science, Yale University, New Haven, CT, United States +1
Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific manifolds, rely on Euclidean approximations, or require costly and numerically fragile geometric operations. This thesis develops a unified framework for Riemannian deep learning from three complementary perspectives: reusable neural modules, manifold-specific network architectures, and the design of underlying geometries. It generalizes batch normalization from Euclidean spaces and individual manifolds to broad classes of Lie groups and gyrogroups, and extends multinomial logistic regression from Euclidean space to SPD manifolds and then to general Riemannian manifolds. It further develops neural networks for several important geometric representations, including an unconstrained model of hyperbolic space, Busemann-based hyperbolic learning, and full-rank correlation matrices. Finally, it introduces adaptive and computationally efficient Riemannian metrics on SPD manifolds, including learnable Log-Euclidean geometries and fast, stable Cholesky-based geometries. The proposed methods are supported by theoretical analysis and validated through numerical experiments and applications in vision, signal processing, graph learning, and genomics.
Ziheng Chen
Doctoral School in Information and Communication Technology, Università di Trento
Hyperbolic neural networks introduce geometric operations that require explicit treatment in relevance propagation. Equivalent geometric realizations can produce different feature attributions, even when local relevance is conserved. We study this problem through Geometric Representation Invariance (GRI), a specialization of Implementation Invariance, and zero-curvature consistency, which requires identity relevance propagation when a geometric module approaches the identity. We propose LRP-radial-all for origin-centered radial modules, treating geometric scaling as modulation and assigning relevance entirely to the signal branch. The rule conserves relevance, is invariant to equivalent radial factorizations, and satisfies zero-curvature consistency, yielding GRI for a specified Poincaré-Lorentz logarithmic-map construction. In contrast, a conservative LRP-half baseline can violate both consistency criteria. Experiments on hyperbolic MNIST, sEEG, and CIFAR-10 classifiers assess attribution fidelity, qualitative explanations, and runtime. LRP-radial-all achieves competitive attribution fidelity across datasets with runtime comparable to Gradient×Input and substantially lower than Integrated Gradients. These findings motivate geometry-aware propagation rules that distinguish relevance conservation from consistency across equivalent computations.