cs.LGOct 4, 2026

Do Neural Networks Learn Structure-Preserving Maps? A Case Study in Latent-to-Hilbert Embeddings

Authors: Muhammad Adnan Shahzad

Organizations: Department of Computer Science and Software Engineering (CSSE) Montreal, QC, Canada

Abstract

We ask whether a neural network can learn a structure-preserving map from a compressed latent space to a Hilbert-space representation. Using an 8-dimensional autoencoder bottleneck on MNIST and nn-qubit product-state targets from PCA-based angle encoding, we report four findings. Although the target angles are generated by a nonlinear sigmoid transformation of the latent projections, the resulting mapping is well approximated by a linear function over the observed latent distribution: linear regression from zz to the true target angles achieves R2=0.98R^{2} = 0.98, while regression to the MLP's recovered angles achieves R2=0.91R^{2} = 0.91. The learned map's primary direction is strongly aligned with the target-induced direction, with cosine similarity 0.9890.989, while remaining nearly orthogonal to the input's principal direction, with cosine similarity 0.0020.002. The map is genuinely rank-4: removing any singular direction degrades inner-product preservation by 2.52.5--5.9×5.9\times despite a singular-value spectrum with two dominant and two small values. The learned subspace does not coincide with the PCA basis used to construct the target, and different random seeds recover the same primary direction but diverge in higher ranks. Finally, kernel ridge regression with an RBF kernel outperforms a tuned MLP (IP error 0.01440.0144 vs.\ 0.01970.0197), suggesting that for approximately linear structure-preserving mappings, classical kernel methods may be a simpler and more effective alternative.

Figures & tables

Explore similar work

Aug 12, 2026cs.LG

Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks

Neural networks can often be trained or fine-tuned through random low-dimensional reparameterization, where a small latent vector is mapped into a full parameter update by a frozen random map. This raises a practical question: how large must the latent search space be to reach a low-loss region? We first express the known accessibility transition in an equivalent conic form, centered for compact convex targets at the statistical dimension of the polar cone. Our main theoretical contribution is an orientation-resolved quadratic master formula that predicts the random-slice residual from both the curvature spectrum and the reference-to-solution displacement profile. It yields a self-consistent isotropic-orientation predictor and, in a conservative radius-only specialization, recovers the earlier Gaussian-width quadratic bound. Building on this analysis, we introduce Random Mapping Networks (RaMaN), which instantiate the predicted latent dimension using structured Hadamard or seed-regenerated Gaussian maps. These constructions avoid the O(dP) storage of dense random maps and reduce optimizer-state memory from O(P) to O(d). We also develop matrix-free curvature approximations and sweep-free dimension selection. Across controlled quadratic and neural-curvature experiments, the orientation-resolved predictor closely tracks measured transition locations and outperforms orientation-agnostic approximations when displacement direction matters. End-to-end experiments further show sharp, protocol-dependent training transitions across image and language models.
May 9, 2026cs.LG

Bilinear autoencoders find interpretable manifolds

Sparse autoencoders have become a standard tool for uncovering interpretable latent representations in neural networks. Yet salient concepts often span manifolds that current linear methods cannot capture without post hoc analysis. This paper uses quadratic latents to close this gap: we implement these with bilinear autoencoders, which decompose activations into low-rank quadratic forms, compose linearly in weight space, and admit input-independent geometric analysis. This qualitative difference in what concepts quadratic latents can detect challenges the standard linear representation hypothesis. Our experiments and visualisations show that multi-dimensional geometries are highly prevalent and that composite latents capture them well, systematically improving reconstruction error in language models. Furthermore, we show that autoencoders with varying geometric priors recover the same input subspace despite their dictionary entries being distinct. Practically, these models serve as an unsupervised tool for manifold discovery, which we demonstrate through an interactive online visualizer for Qwen 3.5. This is a step toward nonlinear but mathematically tractable latent representations whose composition is expressive and interpretable by design.
Aug 1, 2026cs.LG

Kilobyte Models: Neural Networks as a Seed and a Quantized Latent

The cost of storing and transmitting a trained neural network scales with its parameter count, a bottleneck for over-the-air updates, on-device libraries, and other bandwidth-bound deployments. We study an extreme form of model compression in which the deployable artifact is not the weights but a short recipe for regenerating them. Building on Mapping Networks, which express a network's weights as a nonlinear function of a compact trainable latent and a fixed random basis, we observe that only the latent need be stored, because the basis and initialization center are reproducible from an integer seed. A model becomes a seed together with a quantized latent, whose size is set by the latent dimension and bit width rather than the parameter count. We formalize this artifact and introduce a seeded block-wise basis that scales to networks whose projection cannot be held in memory. In our experiments, a mapped model is as accurate as the same network quantized aggressively to a few bits per weight, while taking far fewer bytes to store. Reaching the most aggressive bit widths depends on fine-tuning the latent with quantization in the loop. The results do not depend on the particular random basis, and a structured basis lets the weights be regenerated almost for free even for large networks.