cs.LGJul 12, 2026

Sharp Concentration Bounds for Bundle-Valued Statistics on Manifolds

Authors: Swagatam DasVaclav Snasel

Organizations: Electronics and Communication Sciences Unit Indian Statistical Institute Kolkata 700108, India · VSB Technical University of Ostrava Ostrava, Czech Republic

Abstract

Many geometric statistics and manifold learning pipelines routinely produce observations -- such as tangent vectors or local frames -- whose natural home is a varying family of fibers attached to different points of a base manifold, rather than a single shared vector space. Forming empirical averages requires transporting these observations to a common reference fiber, thereby introducing curvature- and holonomy-driven effects that are absent from classical concentration theory. We develop a non-asymptotic concentration theory for such transported empirical means, deriving finite-sample, dimension-free Hoeffding- and Bernstein-type bounds via sharp Hilbert-space inequalities. When shortest paths to the reference point are non-unique, transport becomes path-dependent and introduces a deterministic holonomy bias; we isolate and quantify this bias through bundle curvature and loop geometry, with sharp closed-form formulas for the tangent bundle of a round sphere. The resulting bias-variance decomposition separates the stochastic fluctuation decaying at the classical n1/2n^{-1/2} rate in sample size nn, from a curvature-driven error floor that no amount of additional data can eliminate; minimax lower bounds confirm both terms are unavoidable. We further establish a robust median-of-means estimator achieving optimal rates under heavy tails and the central limit theorem in the reference fiber. Controlled experiments on the sphere validate all theoretical predictions.

Explore similar work

Apr 20, 2026math.ST

Horospherical Depth and Busemann Median on Hadamard Manifolds

\We introduce the horospherical depth, an intrinsic notion of statistical depth on Hadamard manifolds, and define the Busemann median as the set of its maximizers. The construction exploits the fact that the linear functionals appearing in Tukey's half-space depth are themselves limits of renormalized distance functions; on a Hadamard manifold the same limiting procedure produces Busemann functions, whose sublevel sets are horoballs, the intrinsic replacements for halfspaces. The resulting depth is parametrized by the visual boundary, is isometry-equivariant, and requires neither tangent-space linearization nor a chosen base point. For arbitrary Hadamard manifolds, we prove that the depth regions are nested and geodesically convex, that a centerpoint of depth at least 1/(d+1)1/(d+1) exists, and hence that the Busemann median exists for every Borel probability measure. Under strictly negative sectional curvature and mild regularity assumptions, the depth is strictly quasi-concave and the median is unique. We also establish robustness: the depth is stable under total-variation perturbations, and under contamination escaping to infinity the limiting median depends on the escape direction but not on how far the contaminating mass has moved along the geodesic ray, in contrast with the Fréchet mean. Finally, we establish uniform consistency of the sample depth and convergence of sample depth regions and sample Busemann medians; on symmetric spaces of noncompact type, the argument proceeds through a VC analysis of upper horospherical halfspaces, while on general Hadamard manifolds it follows from a compactness argument under a mild non-atomicity assumption.
Yangdi Jiang, Xiaotian Chang, Cyrus Mostajeran
Jun 4, 2026cs.LG

Efficient Mean Curvature Computation on High-Dimensional Data Manifolds

Estimating local mean curvature at each point of a high-dimensional dataset is a key ingredient of geometry-aware machine learning algorithms, such as the Mean Curvature Boundary Points (MCBP) method. The naive implementation of this computation, based on a local shape operator approximated from k-nearest neighbor patches, involves an explicit construction of a matrix HH whose trace form yields an O(m4)O(m^4) cost per point, rendering the approach intractable for datasets with more than a few dozen features. This paper introduces two complementary contributions that together reduce this cost by several orders of magnitude. The first contribution is an exact algebraic identity. This identity, derived from the orthogonality of the eigenvectors of the covariance matrix and the cyclicity of the trace operator, eliminates HH entirely and reduces the per-point cost to O(m2)O(m^2) after the eigendecomposition. The second contribution addresses the remaining O(m3)O(m^3) bottleneck of the full eigendecomposition. Since the local covariance matrix has rank at most k1mk-1 \ll m, we replace it with a truncated SVD of the k×mk \times m centered data matrix, an O(k2m)O(k^2 m) operation, and derive an analytical approximation for the contribution of the null-space eigenvectors based on the expected value of their outer product under the Haar measure. The resulting estimator has total cost O(k2m+kmp2)O(k^2 m + k m p^2), where p=k1p = k-1. Experiments on real-world datasets confirm speedups of 50 to 300 times relative to the original implementation, with negligible loss when the fast estimator is used to replace the original version. By providing a scalable and data-driven estimate of local curvature, the proposed method establishes curvature as a practical geometric feature for a broad range of machine learning tasks, from classical to modern deep learning pipelines.
Alexandre L. M. Levada
Jun 5, 2026math.ST

A Temporal Spatial Minimax Rate for Smoothly-Varying Distributions in Wasserstein Space

We study the minimax rate of estimating a future value μtn+hμ_{t_n+h} of a curve tμtt\mapstoμ_t in the 22-Wasserstein space P2(Rd)\mathcal{P}_2(\mathbb{R}^d) from finitely many noisy snapshots of its past, under an adiabatic bound tkvε\|\nabla_t^k v\|\le\varepsilon on the kk-th covariant derivative of the velocity field. Our central result is a unified temporal-spatial minimax lower bound: over regular, locally transport-rich subclasses, every estimator incurs W2W_2-risk with MM-exponent γd(k+1)/(k+1+γd)γ_d(k+1)/(k+1+γ_d), γd=min(1/d,1/2)γ_d=\min(1/d,1/2) (MM the total sample size). It follows from a temporal-to-spatial reduction: the smoothness budget defines a reachable W2W_2-ball into which a transport packing is embedded along the time axis, and the information of the entire snapshot experiment is controlled by a Fano argument -- the spatial packing is classical, but its smoothness-admissible temporal embedding and the full-window analysis are new. The bound interpolates a dimension-free extrapolation floor of order εhk+1\varepsilon h^{k+1} -- the irreducible cost of an unobserved future, present even with the exact past -- and the spatial estimation curse MγdM^{-γ_d}, recovering the static distribution-estimation rate as kk\to\infty. We state the lower bound in a design-dependent form -- with a design-weighted effective sample size -- valid for arbitrary observation times, and obtain the closed-form exponent in the dense (equispaced) regime. The matching upper bound is established at k=0k=0 (rate M1/(d+1)M^{-1/(d+1)}, d3d\ge3) and, in a translation submodel, for all kk; for k1k\ge1 a covariant estimator attains the rate conditionally on two estimates (a comparison-geometry bias bound and an optimal-transport map-estimation rate), leaving the unconditional general-kk upper bound as an open problem. Numerical experiments on synthetic curved and flat families corroborate the predicted exponents.
Munsik Kim