cs.LGOct 1, 2026

Neural scaling laws and evolution of learnable activation functions of Kolmogorov-Arnold networks

Authors: Tilen Cadez, Sanghoon Lee, Kyoung-Min Kim

Organizations: Asia Pacific Center for Theoretical Physics, Pohang, Gyeongbuk, 37673, Republic of Korea · Department of Physics, Pohang University of Science and Technology, Pohang, Gyeongbuk 37673, Republic of Korea

Abstract

Kolmogorov-Arnold Networks (KANs) represent a compelling alternative to traditional Multi-Layer Perceptron (MLP)-based neural networks. By employing activation functions as learnable elements, KANs offer superior interpretability, making them suited for scientific domains. In this work, we investigate the neural scaling laws of KANs and the structural evolution of their learnable activation functions under dataset expansion. Specifically, we evaluate the scaling behavior of three KAN variants---BSRBF-KAN, Gottlieb-KAN, and Faster-KAN---across standard image classification benchmarks (MNIST and Fashion-MNIST) and a specialized scientific regression task (magnetic parameter estimation from domain images of moiré magnetic textures). Our results demonstrate that the test loss L{\cal L} exhibits a broken neural scaling law (BNSL) behavior as a function of the dataset size NDN_D. After passing through a random-guess regime, the loss follows architecture- and task-dependent scaling behavior. The loss crosses from a faster- to a slower-scaling branch, L∝ND−α{\cal L}\propto N_D^{-α} and L∝ND−β{\cal L}\propto N_D^{-β} with α>βα>β for image classification tasks. The exponents αα and ββ depend strongly on both the specific network architecture and the dataset-size regime, ranging from 0.4 to 1.5 and from 0.06 to 0.6, respectively. For the magnetic parameter-regression task, the loss follows a single scaling law with its exponent ranging from 1.28 to 2.59. Additionally, we provide a structural analysis of how activation functions refine their complexity as data volume increases, finding that dataset expansion drives a transition from simple linear-like approximations toward stable, interpretable symbolic forms. These findings provide a quantitative roadmap for the efficient application of KANs while managing the trade-off between model expressivity and computational overhead.

Figures & tables

Explore similar work

CardsList
  1. Scale-Parameter Selection in Gaussian Kolmogorov-Arnold Networks

    Apr 23, 2026Amir Noorizadegan, Sifan WangKolmogorov-Arnold NetworksBasis Functions

  2. Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing Systems

    Sep 7, 2026Guangneng Chen, David Garcia Selfa, Pablo Quesada BarriusoKolmogorov-Arnold NetworksHigh-Performance Computing

  3. SparseKAN: Compressing Kolmogorov--Arnold Networks Across Basis Functions, Neurons, and Bits

    Aug 1, 2026Kazi Ahmed Asif Fuad, Lizhong ChenKolmogorov-Arnold NetworksCompressed Model