Tangent Space

Recent momentum

emerging

2 papers in the last 28 days · 0.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

1 new paper

A weekly snapshot of new work published in Tangent Space.

Period ending 2026-09-07

1 new paper

A weekly snapshot of new work published in Tangent Space.

22 papers

Latest in Tangent Space

Sep 14, 2026cs.RO

Trajectory Bundle Method in SE(3) for Black-Box Fixed-Wing Aircraft Trajectory Optimization

Dynamically feasible trajectory optimization for rigid-body systems is naturally formulated on the special Euclidean group SE(3) but is challenging when dynamics are available only as black-box computations without derivatives. This paper formulates the Trajectory Bundle Method (TBM) for motion planning implicitly on SE(3). Bundles are constructed in the Lie algebra and propagated through nonlinear rigid-body dynamics using exponential and logarithmic maps, enabling derivative-free planning of non-Euclidean trajectories. We show that Euclidean TBM interpolation error is bounded quadratically by bundle diameter and extend this result to SE(3), where the bound additionally depends on a local Lipschitz constant of the Log map. Numerical experiments corroborate these bounds. Finally, we demonstrate SE(3) TBM by optimizing an acrobatic, collision-free fixed-wing maneuver through a rotated aperture without explicit models or derivatives of the vehicle dynamics, aerodynamics, or collision model.
Matthew D. Osburn, Cameron K. Peterson, John L. Salmon
Sep 1, 2026stat.ML

Matched Queries for Curvature and Density at Branching Junctions

At a junction, a score field can reveal weighted tangent rays, yet these first-order quantities do not determine how individual branches bend or how their densities change away from the center. Recovering this missing information is necessary for describing local continuation beyond a single point, but finite observations must separate branchwise second-order effects while allowing error in the estimated center. We address this inverse problem using matched score queries at noise scales σσ and λσλσ. For a finite union of C2,αC^{2,α} half-branches in RD\mathbb{R}^D, the normalized score has the expansion Fσ=F0+σG+O(σ1+α)F_σ=F_0+σG+O(σ^{1+α}). Matched subtraction cancels the tangent contribution and exposes GG, which depends linearly on branchwise curvature and log-density slope. Given tangent directions and weights on distinct rays, GG uniquely identifies all sDsD branch parameters, and sDsD scalar component observations are necessary. An O(σ2)O(σ^2) center error introduces DD translation modes, leading to (s+1)D(s+1)D observations under full-rank calibration, except for a translation-invariant full line. We also establish a perturbation bound and a conditional kernel-density-estimation rate. Experiments reproduce the predicted population and N−1/5N^{-1/5} trends and remain full rank up to D=20D=20 with 16 supplied branches. In end-to-end tests for D=3D=3--55, a known-count first-order frontend yields full rank in all 135 population systems and a median relative jet error of 0.132. With strong first-order error, matched responses reduce median parameter error by a factor of 49.4 relative to naive tangent subtraction.
Ziqi Zhao, Qingjian Ni
Aug 7, 2026quant-ph

Readout-Rank Laws for Isotropic Quantum Tangents

Deep parameterized quantum circuits may remain sensitive to a parameter change while the observables retained by a learning model barely respond. We study this separation for a fixed computational-basis measurement. For a pure-state tangent, we compare the quantum Fisher information FQF_Q, the Fisher information FfullF_{\rm full} in the complete bitstring distribution, and the largest variance-normalized response IA\mathcal I_{\mathcal A} available to a diagonal readout space A\mathcal A. If the joint state--tangent frame is Haar random, we prove that the two successive information fractions are independent Beta variables whose means are 1/21/2 and r/(2n−1)r/(2^n-1), where rr is the centered dimension of the readout. Consequently, even the joint span of all computational-basis Pauli strings through any fixed weight kk retain only O(nk2−n)O(n^k2^{-n}) of the full-record information. Exact-statevector experiments across six circuit families show increasing finite-size agreement with this hierarchy in five nonconserving ensembles as the circuit depth grows. A number-conserving family departs strongly from the isotropic prediction even after correcting the support and readout rank, showing that rank alone is insufficient without tangent isotropy.
Marwan Ait Haddou
Aug 3, 2026cs.LG

Feed-Forward Steering in Transformer Residual Dynamics

Attention-only dynamical theories model Transformer residual directions as particles aggregating on a sphere. We extend this framework by incorporating the feed-forward network (FFN) term as a local steering field acting on each token state. The resulting theory predicts that the tangential component of the FFN field is necessary for motion in residual-direction space, that critical residual directions correspond to nonlinear projective equilibria, and that a commutator defect determines when a finite attention--FFN block can be accurately approximated by a parallel, additive flow. Across GPT-2, Pythia, Mistral, and Llama models, the extended theory improves one-step angular prediction relative to an attention-only baseline, with the contribution of the FFN increasing from GPT-2 to Llama-3-8B. Intervention experiments show that retaining only the tangential FFN component preserves most model quality, whereas retaining only the radial component causes performance to collapse. The tangential component also preserves output diversity under aggregation pressure. As a practical application, layers with small commutator defects can be approximately parallelized with only a modest increase in loss, whereas layers with large defects degrade rapidly. These findings support the interpretation of FFN layers as directional steering fields that shape Transformer residual geometry and govern the feasibility of block-level interventions.
Timur Mudarisov, Mikhail Burtsev, Radu State
Jul 26, 2026stat.ME

A Characterization of the Orthocomplement of the Tangent Space of Semiparametric Markov Models

Graphical models are ubiquitous in social and empirical science as they are intuitive and easy to use. These models belong to the broader class of Markov models, defined using solely conditional independence (CI) restrictions. In order to estimate finite-dimensional target parameters in such models efficiently, semi-parametric theory provides a principled framework for constructing regular and asymptotically linear estimators via influence functions (IFs). These estimators are asymptotically normal and root-nn consistent. Characterizing the class of all influence functions for a target parameter is crucial for statistically efficient inference in these models. For models that are Markov relative to directed acyclic graphs (DAGs), the orthogonal complement of the tangent space is known, implying that for any target the class of all influence functions can be derived once an influence function is obtained. On the other hand, for Markov models not equivalent to a DAG model -- such as ordinary Markov models associated with undirected graphs, chain graphs, or acyclic directed mixed graphs -- the orthogonal complement has not been characterized, impeding semi-parametric inference in these models. We derive closed form expressions for the orthogonal complement of the tangent space for general Markov models and illustrate our results by characterizing the class of influence functions for the conditional mean parameter in several graphical models.
Trung Phung, Ilya Shpitser
Jul 16, 2026cs.LG

Learning in Infinitesimal Non-Compositional Sketches

This paper develops a categorical framework -- Learning in Infinitesimal Non-Compositional Sketches (LINCS) -- as the repair of non-compositionality: failures of diagrams to factor through quotient sketches lifted to the tangent category setting. Machine learning problems are specified as sketches: graphs with commutativity conditions D\mathcal D, limit cones L\mathcal L, and colimit cocones K\mathcal K, generalizing the usual scalarization of loss functions or vector space assumptions. Non-compositionality is defined purely as failure of a universal factorization problem, not as arithmetic error between the desired and actual predictions. Given a learning sketch S=(S,D,L,K)\mathbb S=(S,\mathcal D,\mathcal L,\mathcal K), whose underlying graph is SS, and a model D:J→CD:J \rightarrow C, the base defect is the obstruction to factorization \mboxObs(\mboxFactS(D))\mbox{Obs}(\mbox{Fact}_{\mathbb S}(D)). The tangent lift applies the tangent functor TT to obtain TD:J→CTD:J \rightarrow C, and LINCS is defined as the obstruction \mboxObs(\mboxFactS(TD))\mbox{Obs}(\mbox{Fact}_{\mathbb S}(TD)) -- asking whether infinitesimal perturbations preserve the compositionality constraints.The paper also introduces Tangent Learning Sketches, which are sketches equipped with Cockett-Cruttwell tangent structure. The paper defines the INC endofunctor, which iterates the tangent lift, producing a tower D,TD,T2D,⋯D,TD,T^2D, \cdots of factorization problems. ML is thereby formulated as the search for a coalgebraic fixed point where successive tangent unfoldings stabilize (νT\mboxINCνT_{\mbox{INC}}). Using the Aczel--Mendler theorem, we prove existence of a final INC coalgebra whenever T\mboxINCT_{\mbox{INC}} admits a set-based class realization that creates its final carrier. A detailed experimental evaluation of LINCS is underway in a number of concrete ML settings, including deep learning, large language models, and reinforcement learning, and is described in companion papers.
Sridhar Mahadevan
Jul 7, 2026math.AG

Tangent classes of matroids and wonderful compactifications

For every loopless matroid MM and every Feichtner--Yuzvinsky building set G\mathcal{G} containing the top flat, we construct an integral tangent class TM,GZ∈KZ(M,G)T_{M,\mathcal{G}}^{\mathbb{Z}}\in K_{\mathbb{Z}}(M,\mathcal{G}); in the realizable case it specializes to the class of the tangent bundle of the corresponding wonderful compactification, it recovers the Hilbert series of the Chow ring through Hirzebruch--Riemann--Roch, and it satisfies the expected Chern-alpha lower bounds. This reproduces the tangent class and its key properties studied by the first author in arXiv:2606.22650. The main body of this paper was produced autonomously, without human mathematical guidance, by Danus, an AI mathematical reasoning agent. Danus solved the problem before arXiv:2606.22650 was publicly available, demonstrating the potential of AI agents in mathematical research. We reproduce its output faithfully, adding only editorial comments; the experiment is documented in Appendix B.
Ronnie Cheng, Shurui Liu, Guoxiong Gao
Jun 23, 2026math.CT

Infinitesimal Causality

This paper introduces a categorical account of infinitesimal causality in Frobenius Markov categories equipped with tangent-bundle semantics. IDC captures the infinitesimal layer in which interventions act as tangent deformations of copy/discard structure. Two distinct Frobenius structures interact: (1) the categorical Frobenius algebra on classical variables encoding copying, comparing, and discarding; and (2) the geometric Frobenius integrability condition, namely involutive closure of the intervention distribution, distinct from the algebraic Frobenius structure. Categorical causal sufficiency is defined as the compatibility of these two notions. A key observation is that, for structural causal models, infinitesimal causality is most naturally formulated in the slice of deterministic mechanisms over exogenous variables, with visible stochastic kernels obtained only after pushforward. Interventions are tangent vectors that deform the Frobenius copy/discard operations; their Lie brackets measure whether this deformation preserves classical information-flow structure. Pearl's do-calculus is used as a guiding example of intervention identities: ignoring irrelevant interventions corresponds to counit invariance, action/observation exchange to coproduct compatibility with pushforward, and independence to involutive bracket closure of the visible intervention distribution.
Sridhar Mahadevan
Jun 10, 2026cs.LG

Scalable anomaly detection via a univariate Christoffel function

Anomaly detection plays a critical role in identifying unusual patterns across domains such as fraud detection, network intrusion, and system fault diagnosis. Recently, Christoffel function-based methods, rooted in polynomial optimization, have emerged as promising alternatives to deep learning due to their strong mathematical foundations and computational frugality. However, their practical applicability is hindered by the need to invert a matrix whose size grows exponentially with the data dimension, rendering the method intractable even for moderate-dimensional datasets. This paper addresses the dimensionality limitations of Christoffel function-based anomaly detection while preserving its key theoretical properties, i.e., the on-off support dichotomy behavior and the accurate support shape capture. We introduce UCF, a univariate Christoffel function which is based on the squared distance between the query point and the support points. Extensive experiments on the ADBench benchmark demonstrate that UCF consistently outperforms 14 state-of-the-art baselines in terms of Average Precision. By resolving the scalability bottleneck of the Christoffel Function, this work expands the toolkit of anomaly detection methods with a robust, theoretically grounded, and universally applicable approach.
Florian Grivet, Didier Henrion, Jean-Bernard Lasserre +1
Jun 10, 2026cs.LG

Tree-Structured Orthonormal Decomposition of the Aitchison Simplex

Compositional data -- vectors encoding relative proportions -- arise across scientific domains, including ecology, geochemistry, and genomics. The features in these data often come with known hierarchical structure (e.g., taxonomies, phylogenies, ontologies), yet existing methods either ignore this structure, discard the intrinsic Aitchison geometry, are designed for binary trees, or yield incomplete coordinate systems. We describe PolyILR, a canonical orthonormal decomposition of the Aitchison tangent space aligned with any tree topology. Our construction defines a weighted local geometry at each internal node capturing full branching structure, then lifts these to a global orthonormal basis where every coordinate corresponds to a specific tree location. On microbiome and single-cell benchmarks, PolyILR yields stable, interpretable features and enables inference at multiscale tree resolution. We also establish a novel theoretical connection to softmax classifiers, suggesting possible applications to probabilistic modeling.
Daisuke Yamada, Qijun Zhang, Travis Pence +3
Jun 8, 2026physics.comp-ph

Input-schema identifiability limits in physics-informed surrogates for mechanics-governed flow

Physics-informed and data-driven surrogates are increasingly used to approximate mechanics-governed flow fields, but the target quantities assigned to such models are not always identifiable from the input variables available at prediction time. We introduce an input-schema identifiability certificate for computational surrogates. Starting from a reduced physical model, the certificate decomposes a target field into components that are measurable from geometry, components that require boundary-condition information, and components identifiable only up to a symmetry quotient. This yields a pre-training audit: it predicts which oracle-channel interventions should reduce error, which should fail, and which ambiguity cannot be removed by changing the architecture, loss, optimizer, or sample size. We instantiate the framework for incompressible tubular flow using a Cosserat-rod reduction, where lumen velocity separates into a mesh-measurable tangent direction, a boundary-condition-dependent magnitude, and a signed-orientation ambiguity. Controlled experiments on patient-specific aortic CFD geometries, analytic Womersley flows, and an advection-diffusion transfer problem confirm the predicted pattern: supplying signed direction collapses angular error to the oracle regime, whereas supplying magnitude without orientation leaves the predicted sign ambiguity and yields 16-33 percent per-node sign flips. The results provide a mechanics-based diagnostic for deciding whether a surrogate modelling task is physically identifiable before training, and expose failure modes that aggregate error metrics can hide.
Daniel Cieslak, Andrzej Czyzewski
Jun 7, 2026cs.DC

Aperon Technical Report: Hierarchical No-Pointer Tangent-Local Search for High-Dimensional Approximate Nearest Neighbors

We present HNTL (Hierarchical No-pointer Tangent-Local), the core vector indexing and candidate generation framework of the Aperon vector memory system. Proximity graphs (e.g., HNSW) incur a heavy pointer tax in memory overhead and induce irregular memory accesses that stall CPU pipelines. HNTL resolves this by partitioning the high-dimensional space into local, coherent grains, representing vectors as low-dimensional coordinates on local tangent spaces, and scanning them sequentially using a pointerless Block-SoA (Structure-of-Arrays) layout. On anisotropic manifold data (d=768, N=10,000), local PCA captures 96.3% of the variance, allowing HNTL to achieve a final Rerank Recall@10 of 1.0000 with a candidate pool size of only C=20 vectors. Hardware profiling via Apple kperf CPU Performance Monitoring Unit (PMU) counters demonstrates a 3.61x speedup (4.137 ns/vector vs. 14.951 ns/vector) for our NEON auto-vectorized C++ Block-SoA scan engine over standard pointer-chasing graph traversals, driven by a 3.59x IPC (Instructions Per Cycle) and near-zero L1/L2 data cache misses.
Yong Fu
Jun 1, 2026math.NA

Spectral Audit of In-Context Operator Networks

Existing evaluations of neural operators and in-context operator learning rely primarily on prediction error, but accurate output prediction does not guarantee the correct local dynamical structure. A model may match solutions while exhibiting incorrect sensitivities, distorted frequency response, spurious mode coupling, or unstable tangent behavior. We introduce a Jacobian-based spectral audit for in-context operator learning. For a fixed prompt, we differentiate the network output with respect to the query function and view the resulting Jacobian as a learned tangent operator. Projecting it onto Fourier modes, we obtain a local spectral characterization of the inferred operator, including frequency-dependent gains, phase structure, and cross-mode coupling. The audit complements standard prediction metrics by testing whether the model reproduces local mechanisms of the underlying PDE operator rather than only outputs. Across benchmarks, the audit reveals distinct operator-level phenomena, including phase transport, viscosity-dependent damping, nonlinear mode coupling, and reaction--diffusion stability structure. It also detects failures partially hidden by prediction-error metrics, including high-frequency degradation, incorrect phase recovery, and prompt--operator inconsistencies. Corrupted or internally inconsistent prompts lead to degraded tangent-operator structure even when pointwise predictions remain partially accurate. Our results suggest that prediction accuracy and local operator fidelity are distinct properties of learned neural operators. Our framework also provides a diagnostic for stability, sensitivity, and operator consistency.
Zhiwei Gao, Liu Yang, George Em Karniadakis
Jun 1, 2026cs.RO

The Lie We Tell: Correcting the Euclidean Fallacy in Vision Language Action Policies via Score Matching on Tangent Space

Diffusion-based Vision-Language-Action policies achieve remarkable success in robotic manipulation, yet commit a fundamental geometric error we term the Euclidean Fallacy\textbf{Euclidean Fallacy}: representing SE(3) poses as flat R12\mathbb{R}^{12} vectors. This approximation induces (1) manifold drift violating SO(3) constraints, (2) broken equivariance under coordinate transformations, and (3) non-geodesic trajectories with excessive kinematic cost. We introduce Lie Diffuser Actor (LDA)\textbf{Lie Diffuser Actor (LDA)}, a diffusion framework operating intrinsically on SE(3). Our method injects noise through left-invariant SDEs, predicts scores in the tangent space, and retracts samples via the exponential map. This formulation eliminates manifold drift by construction while guaranteeing coordinate-frame equivariance and geodesic optimality. On CALVIN ABC→\rightarrowD, LDA improves average task length from 3.273.27 to 3.513.51 (+7.3%+7.3\%). We further validate our method on real robot and the results show that our methodology outperforms the baseline on majority tasks.
Bing-Cheng Chuang, I-Hsuan Chu, Bor-Jiun Lin +3
May 28, 2026cs.CV

Geodesics with Unified Tangent-constrained Priors and Curvature Regularization

Curvature-penalized geodesic models have proven their effectiveness in image segmentation by computing globally optimal curves. Unfortunately, these models remain susceptible to shortcuts when delineating objects with complex shapes and image intensity distributions, as they lack mechanisms to enforce shape-aware tangent constraints. To address this limitation, we propose a unified geodesic framework that integrates tangent-constrained priors with curvature penalization. The key idea is to formulate tangent admissibility directly within the orientation-lifted space, where path tangents are restricted to spatially varying angular sectors derived from intrinsic shape representatives (ISR) such as skeletons or interior landmarks. This formulation gives rise to a family of tangent-constrained Finslerian metrics, extending the classical curvature-penalized geodesic models while enforcing mandatory tangent constraints. The resulting Hamilton-Jacobi-Bellman (HJB) partial differential equations (PDEs) admit efficient numerical solutions via variants of the fast marching method, preserving the single-pass computational complexity. Experiments on synthetic, natural, and medical images demonstrate that the proposed geodesic framework indeed improves robustness against weak boundaries and topological shortcuts, yielding segmentation results with enhanced shape fidelity compared to existing geodesic models.
Chong Di, Li Liu, Jinglin Zhang +3
May 13, 2026cs.RO

Identification of Non-Transversal Bifurcations of Linkages

The local analysis is an established approach to the study of singularities and mobility of linkages. Key result of such analyses is a local picture of the finite motion through a configuration. This reveals the finite mobility at that point and the tangents to smooth motion curves. It does, however, not immediately allow to distinguish between motion branches that do not intersect transversally (which is a rather uncommon situation that has only recently been discussed in the literature). The mathematical framework for such a local analysis is the kinematic tangent cone. It is shown in this paper that the constructive definition of the kinematic tangent cone already involves all information necessary to separate different motion branches. A computational method is derived by amending the algorithmic framework reported in previous publications.
Andreas Mueller, P. C. López Custodio, J. S. Dai
May 5, 2026cs.NE

Symmetry, Defects, and Diffusion in Continuous-memory Recurrent Networks

Continuous-memory recurrent networks must preserve phase, position, or orientation despite model imperfections and state noise. We develop a geometric framework that separates three questions: how many memory coordinates are neutrally transported, how deterministic perturbations alter their finite-horizon stability, and how ambient noise is decoded along them. Exact per-input equivariance transports analytical group tangents pathwise and, on a compact nondegenerate orbit stratum, yields at least q=dim⁡(G/H)q=\dim(G/H) zero group-tangent Lyapunov exponents under stationary ergodic driving. For imperfect dynamics, a four-block tangent/normal decomposition gives local and finite-horizon bounds on tangent growth and subspace rotation, distinguishing first-order direct damage from second-order leakage through contracting normal directions. For noisy dynamics, a specified decoder maps ambient covariance QQ to coordinate covariance ZQZ⊤ZQZ^\top; under isotropic noise and fixed tangent energy, least-squares decoding and scaled-isometric action geometry minimize local diffusion. Local and finite-horizon evaluations include cases both within and outside the sufficient conditions. In a fresh twenty-seed T2T^2 replication, a decoder-covariance objective improves noisy horizon-256 memory in every pair while meeting a prespecified clean-error equivalence margin. Direct noise training also improves noisy memory but incurs a clean-error tradeoff. In coupled T4/T8T^4/T^8 integrators, the advantage of anisotropic-covariance over isotropic regularization reverses when evaluation noise becomes isotropic. These results connect continuous symmetry to measurable limits and design choices for recurrent memory under explicit dynamical, decoder, and noise assumptions.
Hanson Hanxuan Mo
Apr 24, 2026math.GR

Closed Form Relations and Higher-Order Approximations of First and Second Derivatives of the Tangent Operator on SE(3)

The Lie group SE(3) of isometric orientation preserving transformation is used for modeling multibody systems, robots, and Cosserat continua. The use of these models in numerical simulation and optimization schemes necessitates the exponential map, its right-trivialized differential (often referred to as tangent operator), as well as higher derivatives in closed form. The 6×66\times 6 matrix representation of the differential, dexpX:se(3)→se(3)\mathbf{dexp}_{\mathbf{X}}:se\left( 3\right) \rightarrow se\left( 3\right) , and its first derivative were reported using a 3×33\times 3 block partitioning. In this paper, the differential, its first and second derivative, as well as the Jacobian and Hessian of the evaluation maps, dexpXZ\mathbf{dexp}_{\mathbf{X}}\mathbf{Z} and dexpXT\mathbf{dexp}_{\mathbf{X}}^{T}% \mathbf{Z}, are reported avoiding the block partitioning. For all of them, higher-order approximations are derived. Besides the compactness, the advantage of the presented closed form relations is their numerical robustness when combined with the local approximation. The formulations are demonstrated for computation of the deformation field and the strain rates of an elastic Cosserat-Simo-Reissner rod.
Andreas Mueller
Apr 17, 2026cs.LG

Geometric regularization of autoencoders via observed stochastic dynamics

Stochastic dynamical systems with slow or metastable behavior evolve, on long time scales, on an unknown low-dimensional manifold in high-dimensional ambient space. Building a reduced simulator from short-burst ambient ensembles is a long-standing problem: local-chart methods like ATLAS suffer from exponential landmark scaling and per-step reprojection, while autoencoder alternatives leave tangent-bundle geometry poorly constrained, and the errors propagate into the learned drift and diffusion. We observe that the ambient covariance~ΛΛ already encodes coordinate-invariant tangent-space information, its range spanning the tangent bundle. Using this, we construct a tangent-bundle penalty and an inverse-consistency penalty for a three-stage pipeline (chart learning, latent drift, latent diffusion) that learns a single nonlinear chart and the latent SDE. The penalties induce a function-space metric, the ρρ-metric, strictly weaker than the Sobolev H1H^1 norm yet achieving the same chart-quality generalization rate up to logarithmic factors. For the drift, we derive an encoder-pullback target via Itô's formula on the learned encoder and prove a bias decomposition showing the standard decoder-side formula carries systematic error for any imperfect chart. Under a W2,∞W^{2,\infty} chart-convergence assumption, chart-level error propagates controllably to weak convergence of the ambient dynamics and to convergence of radial mean first-passage times. Experiments on four surfaces embedded in up to 201201 ambient dimensions reduce radial MFPT error by 5050--70%70\% under rotation dynamics and achieve the lowest inter-well MFPT error on most surface--transition pairs under metastable Müller--Brown Langevin dynamics, while reducing end-to-end ambient coefficient errors by up to an order of magnitude relative to an unregularized autoencoder.
Sean Hill, Felix X. -F. Ye
Oct 2, 2025cs.LG

Robust Tangent Space Estimation via Laplacian Eigenvector Gradient Orthogonalization

Estimating the tangent spaces of a data manifold is a fundamental problem in geometric data analysis. The standard approach, Local Principal Component Analysis (LPCA), struggles in high-noise setting due to a critical trade-off in choosing the neighborhood size. Selecting an optimal size requires prior knowledge of the geometric and noise characteristics of the data that are often unavailable. In this paper, we propose a spectral method, Laplacian Eigenvector Gradient Orthogonalization (LEGO), that utilizes the global structure of the data to guide local tangent space estimation. Instead of relying solely on local neighborhoods, LEGO estimates the tangent space at each data point by orthogonalizing the gradients of low-frequency eigenvectors of the graph Laplacian. We provide two theoretical justifications of our method. First, a differential geometric analysis on the tubular neighborhood of a manifold shows that gradients of the low-frequency Neumann eigenfunctions of the tube align closely with the manifold's tangent bundle, while an eigenfunction with high gradient in directions orthogonal to the manifold lie deeper in the spectrum. Second, a random matrix theoretic analysis also demonstrates that low-frequency eigenvectors are robust to sub-Gaussian noise. These results allow us to derive the asymptotic scaling and stability of the estimated eigenvector gradients. Numerical experiments demonstrate that LEGO yields tangent space estimates that are significantly more robust to noise than those from LPCA, resulting in marked improvements in downstream tasks such as manifold learning, boundary detection, and local intrinsic dimension estimation.
Dhruv Kohli, Sawyer J. Robertson, Gal Mishne +1
Jul 21, 2025cs.LG

Inexact calculus of variations on the hyperspherical tangent bundle with connections to the attention mechanism

We offer a theoretical mathematical background through Lagrangian optimization on the unit hyperspherical manifold and its tangential structure. Our methods can be categorized as inexact since our methods are projection-based and since we will perturb the functional optimization with epsilon-type quantities. We draw connections to the attention mechanism and the Transformer since it exists as a flow map in the tangent fiber for each token along the high-dimensional unit sphere. Our motivation for this work is primarily twofold: we study the attention mechanism under its flow map and its relations to traditional calculus of variations and Lagrangian optimization; and we study a range of calculus of variations on the unit hypersphere that appeal to a broader mathematical lens in approximating, variational contexts.
Andrew Gracyk
May 26, 2023cs.LG

Generalizing Adam to Manifolds for Efficiently Training Transformers

One of the primary reasons behind the success of neural networks has been the emergence of an array of new, highly-successful optimizers, perhaps most importantly the Adam optimizer. It is widely used for training neural networks, yet notoriously hard to interpret. Lacking a clear physical intuition, Adam is difficult to generalize to manifolds. Some attempts have been made to directly apply parts of the Adam algorithm to manifolds or to find an underlying structure, but a full generalization has remained elusive. In this work a new approach is presented that leverages the special structure of the manifolds which are relevant for optimization of neural networks, such as the Stiefel manifold, the symplectic Stiefel manifold and the Grassmann manifold: all of these are homogeneous spaces and as such admit a global tangent space representation. This is a common vector space, often called the Lie subspace, that makes the generalization of all steps in the Adam optimizer (as well as other optimizers) possible. It is thus possible to extend the Adam optimizer to manifolds without a projection step, something that was not possible before. The resulting algorithm is then applied to train transformers and a symplectic autoencoder for which orthogonality constraints are enforced up to machine precision and we conclusively demonstrate the advantage of the proposed optimizer over existing methods.
Benedikt Brantner