Modular Arithmetic

Momentum

0 papers in the last four weeks, with none the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 14

All topics
CardsList
  1. Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers

    Aug 7, 2026Ali Janati, Kaoutar El Maghraoui, Anass BelfatmiTransformerMuon Optimizer

  2. Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations

    Jul 15, 2026Chon-Fai Kam, Xavier Cadet, Miloud Bessafi +1Feedforward Neural NetworksNeural Network Expressivity

  3. Multiplication Beyond Groups: Stratified Fourier Mechanisms in Transformer Circuits

    Jul 8, 2026Zitong Andrew Chen, Junaid Hasan, Akhil Srinivasan +2Transformer InterpretabilityTransformer Attention

  4. Prime Fourier Embeddings: A Principled Basis for Modular Arithmetic

    Jun 22, 2026Hyunsang Hwang, Suhyun Bae, Donghun LeeFourier Feature EmbeddingsRepresentation Learning

  5. Beyond Neural Collapse: Task-Intrinsic Geometry Governs Neural Representations in Modular Arithmetic

    Jun 8, 2026Hu Tan, Kuo Gai, Shihua ZhangNeural Representation GeometryClassification

  6. BRo-JEPA: Learning Modular Arithmetic in Latent Space

    May 31, 2026Divyansh Jha, Yuanfang Xie, Varan Mehra +1Zero-Shot LearningJoint-Embedding Predictive Architecture

  7. Unveiling Memorization-Generalization Coexistence: A Case Study on Arithmetic Tasks with Label Noise

    May 18, 2026Linyu Liu, Pinyan LuNeural Network GeneralizationNeural Network Memorization

  8. Learning Large-Scale Modular Addition with an Auxiliary Modulus

    May 8, 2026Hanato Kikuchi, Ryosuke Masuya, Kazuhiko Kawamoto +1Boolean Function LearningDistribution Shift

  9. Universal Quantum Transformer

    Apr 29, 2026Sungyong Chung, Alireza TalebpourVariational Quantum CircuitsModular Arithmetic

  10. Convergent Evolution: How Different Language Models Learn Similar Number Representations

    Apr 22, 2026Deqing Fu, Tianyi Zhou, Mikhail Belkin +2Fourier Feature EmbeddingsFrequency-Domain Feature Learning