Organizations: School of Physical Science and Technology, Soochow University, 333 Ganjiang Road, Suzhou 215006, P.R. China · Institute for Advanced Study, Soochow University, 333 Ganjiang Road, Suzhou 215006, P.R. China · Mandelstam Institute for Theoretical Physics, School of Physics and NITheCS, University of the Witwatersrand, Johannesburg 2050, South Africa · NSF AI Institute for Artificial Intelligence and Fundamental Interactions (IAIFI) and Department of Physics, Northeastern University, Boston, MA 02115, USA
Based on a transformer based sequence-to-sequence architecture combined with a dynamic batching algorithm, this work introduces a machine learning framework for automatically simplifying complex expressions involving multiple elliptic Gamma functions, including the q-θ function and the elliptic Gamma function. The model learns to apply algebraic identities, particularly the SL(2,Z) and SL(3,Z) modular transformations, to reduce heavily scrambled expressions to their canonical forms. Experimental results show that the model achieves over 99% accuracy on in-distribution tests and maintains robust performance (exceeding 90% accuracy) under significant extrapolation, such as with deeper scrambling depths. This demonstrates that the model has internalized the underlying algebraic rules of modular transformations rather than merely memorizing training patterns. Our work presents the first successful application of machine learning to perform symbolic simplification using modular identities, offering a new automated tool for computations with special functions in quantum field theory and the string theory.
Inspired by long-standing open problems in algebraic combinatorics, we show that modern machine learning can meaningfully contribute to verifiable mathematical discoveries. In particular, we focus on the construction of simple mathematical functions under exact distributional constraints, a setting we formalize as Simple Learning Under Rigid Proportions (SLURP). We tackle this problem by introducing two methods: MapSeek-Functional, which models the desired function alternating pseudo-labeling and supervised training steps; and MapSeek-Symbolic, designed to directly produce symbolic formulas. We successfully apply both methods to a research problem in algebraic combinatorics, discovering a new combinatorial interpretation of the q,t-Narayana polynomials arising from representation theory. To our knowledge, this is the first such interpretation based on noncrossing partitions. Using one discovered statistic, we find a combinatorial proof of the symmetry of these polynomials in a previously unsolved case. To streamline verification and reproducibility, we release all code, including a formalization of all the mathematical discoveries of this paper in Lean 4.
Eugenio Cainelli, Lorenzo Luccioli, Alessandro Iraci +2
University of Bologna · Pegaso University · University of Pisa
Learning parity functions, more general modular addition, is a challenging machine learning task due to its input sensitivity. A recent study substantially scaled modular addition learning in both the number of summands and the modulus. Its key idea is to increase zeros in training sequences, reducing the effective number of summands and thus controlling training difficulty; however, this induces covariate shift between training and test input distributions. This study theoretically and empirically analyzes this side effect and proposes a covariate-shift-free method for modular addition. Specifically, we introduce an auxiliary modulus Kq during training, which reduces wrap-around frequency and problem difficulty while preserving the same input distribution across training and testing. Experiments show strong scalability and sample efficiency: even for large input length N, large modulus q, and small datasets -- where the sparse method fails to learn -- our method achieves equal or better match accuracy and relaxed τ-accuracy. For example, at N=64 and q=974269, our method trained on 100K samples achieves 97.0%τ-accuracy at τ=0.05, while the sparse method achieves only 9.5% with the same data size and 93.9% even when extended to 1M samples.
Can neural networks learn abstract algebraic rules, or do they merely memorize training patterns? We investigate this using MNIST digits as states and modular arithmetic operations as actions in a JEPA-style latent world model. Standard supervised baselines and JEPA models with additive operation embeddings fit seen operations but fail to extrapolate reliably to unseen ones. To bridge this gap, we introduce a block-rotation predictor that imposes the circular structure of modulo-10 arithmetic in latent space. This enables strong zero-shot generalization, with the best ResNet-based JEPA block-rotation model achieving 99.46% zero-shot and 99.46% rollout accuracy. Our results suggest that latent world models can learn symbolic transformation rules when architecture matches the structure of the problem. Our code can be accessed here.
Divyansh Jha, Yuanfang Xie, Varan Mehra +1
1Georgia Institute of Technology · 2NYU Langone Health