SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions
Organizations: University of Information Technology, Ho Chi Minh City, Vietnam · Vietnam National University Ho Chi Minh City, Ho Chi Minh City, Vietnam
Abstract
In recent years KolmogorovArnold Networks KANs have attracted increasing attention due to their effectiveness in machine learning and scientific computing offering a new paradigm for neural network design In this paper we present SechKAN a novel KAN based on hyperbolic secant sech functions The hyperbolic secant basis is adopted for its smooth bellshaped form localized responses and wellbehaved gradients We employ a 1D linear projection to reduce the number of parameters allowing SechKAN to maintain a model size comparable to that of multilayer perceptrons MLPs Experimental results show the effectiveness of SechKAN on function fitting PDE surrogate modeling and image classification benchmarks including MNIST FashionMNIST CIFAR10 and CIFAR100 On function fitting SechKAN achieves performance comparable to both MLPs and representative KAN variants On PDE surrogate modeling it outperforms MLPs and achieves competitive or better performance than representative KAN variants On image classification benchmarks SechKAN achieves the best performance among the evaluated KAN variants while remaining competitive with MLPs using a comparable number of parameters However SechKAN still incurs higher computational cost than MLPs and some KAN variants Our source code is publicly available at https://github.com/hoangthangta/All-KAN.
Figures & tables
| Variant | Inner Function | Base Function |
|---|---|---|
| EfficientKAN [ 5 ] | spline coefficient | |
| FastKAN [ 24 ] | ||
| center, function width or spread |
| Basis Function (KAN Variant) | Basis Characteristics | Parameter Reduction |
|---|---|---|
| B-spline (EfficientKAN [ 5 ] ) | Piecewise polynomial, compact support, piecewise smooth, efficient B-spline implementation | – |
| Gaussian RBF (FastKAN [ 24 ] ) | Infinitely smooth, Gaussian decay, smooth bounded gradients, analytic closed form | – |
| RSWAF (FasterKAN [ 10 ] ) | Infinitely smooth, tanh-based localization, smooth bounded gradients, analytic closed form | – |
| ReLU-derived basis (ReLU-KAN [ 29 ] ) | Piecewise quartic polynomial, compact support, piecewise smooth, analytic closed form | – |
| Hyperbolic secant or sech (SechKAN, Ours) | Infinitely smooth, exponential decay, smooth bounded gradients, analytic closed form | 1D linear projection |
| B-spline and Gaussian RBF (PRKAN [ 35 ] ) | Original B-spline and Gaussian RBF basis with dimensionality reduction | Attention, Conv, Conv+Pool, Dim-Sum, Feature Weight Vector |
| Functions | Other KANs | SechKAN |
|---|---|---|
| [1,1] | [1,1] | |
| [1,1] | [1,1] | |
| [2,4,4,1] | [2,14,14,1] | |
| [3,4,4,1] | [3,16,16,1] | |
| [4,8,8,1] | [4,32,32,1] | |
| [3,4,4,1] | [3,16,16,1] |
| Model | Function | Used Params | Loss (MSE) | Time (s) |
|---|---|---|---|---|
| BSRBF-KAN | Avg | 648 | 6.89 | |
| EfficientKAN | Avg | 648 | 6.69 | |
| FastKAN | Avg | 658 | 5.65 | |
| ReLU-KAN | Avg | 643 | 5.73 | |
| SechKAN | Avg | 651 | 6.74 |
| Comparison | Mean Difference | -value | n Functions |
|---|---|---|---|
| SechKAN vs FastKAN | -0.025807 | 0.152600 | 10 |
| SechKAN vs ReLU-KAN | -0.021036 | 0.110100 | 10 |
| SechKAN vs BSRBF-KAN | -0.014759 | 0.152600 | 10 |
| SechKAN vs EfficientKAN | -0.001836 | 0.747600 | 10 |
| Function | Comparison | Mean Difference | -value | n Seeds |
|---|---|---|---|---|
| SechKAN vs EfficientKAN | 0.001945 | 0.1375 | 10 | |
| SechKAN vs FastKAN | -0.095961 | 0.0015 | 10 | |
| SechKAN vs BSRBF-KAN | -0.028931 | 0.0015 | 10 | |
| SechKAN vs ReLU-KAN | -0.063311 | 0.0015 | 10 | |
| SechKAN vs EfficientKAN | -0.003059 | 0.2158 | 10 | |
| SechKAN vs FastKAN | -0.053119 | 0.0015 | 10 |
| Model | Val. Rel. L2 | Test Rel. L2 | Runtime (s) | Used Params |
|---|---|---|---|---|
| BSRBF-KAN | 0.09664 0.00728 | 0.30327 0.01853 | 422.67 3.29 | 39744 |
| EfficientKAN | 0.08496 0.00345 | 0.26524 0.01692 | 388.91 2.15 | 39630 |
| FastKAN | 0.30503 0.01253 | 0.72194 0.02485 | 86.05 1.42 | 38684 |
| FasterKAN | 0.26852 0.06415 | 0.55705 0.13917 | 179.68 51.91 | 39152 |
| MLP | 0.69984 0.00096 | 0.72133 0.00361 | 130.69 21.11 | 39396 |
| ReLU-KAN | 0.78003 0.01123 | 0.89849 0.03682 | 139.67 39.35 | 39782 |
| Model | Val. Rel. L2 | Test Rel. L2 | Runtime (s) | Used Params |
|---|---|---|---|---|
| BSRBF-KAN | 0.15825 0.01704 | 0.41038 0.05549 | 215.60 119.96 | 40320 |
| EfficientKAN | 0.20637 0.01049 | 0.57372 0.03928 | 143.23 40.57 | 40230 |
| FastKAN | 0.21505 0.01030 | 0.59755 0.04715 | 116.11 40.31 | 39252 |
| FasterKAN | 0.22423 0.02903 | 0.39949 0.03972 | 178.42 65.66 | 39688 |
| MLP | 0.41146 0.00383 | 0.46458 0.00813 | 108.28 69.98 | 38016 |
| ReLU-KAN | 0.41802 0.01192 | 0.61521 0.00488 | 700.64 92.31 | 40311 |
| Dataset | Model | Network structure | Used Params | MFLOPs | Peak Mem. (MB) |
|---|---|---|---|---|---|
| MNIST Fashion-MNIST | SechKAN | (784, 64, 10) | 52,608 | 0.096 | 23.84 |
| MLP | (784, 64, 10) | 50,816 | 0.051 | 17.42 | |
| CNN | 2 Conv layers + 2 MLP layers | 52,138 | 0.378 | 19.87 | |
| SechKAN-CNN | 2 Conv layers + 2 SechKAN layers | 52,160 | 0.418 | 21.09 | |
| CIFAR-10 | SechKAN | (3072, 64, 10) | 203,632 | 0.501 | 31.84 |
| MLP | (3072, 64, 10) | 197,248 | 0.198 | 20.21 |
| Dataset | Network | Train. Acc. | Val. Acc. | Val. F1 | Runtime (s) |
|---|---|---|---|---|---|
| MNIST | MLP | ||||
| SechKAN | |||||
| SechKAN-CNN | |||||
| CNN | |||||
| Fashion-MNIST | MLP | ||||
| SechKAN |
| Model | Basis | 1D Projection | Loss | Params | Runtime (s) |
|---|---|---|---|---|---|
| EfficientKAN | B-spline | 648 | 7.02 2.43 | ||
| Modified EfficientKAN | B-spline | 647 | 7.28 2.52 | ||
| Modified SechKAN | sech | 653 | 6.01 1.84 | ||
| SechKAN (Proposed) | sech | 660 | 6.22 1.85 |
| Dataset | Model | Basis | 1D Projection | Val. Acc. | Runtime (s) |
|---|---|---|---|---|---|
| MNIST | EfficientKAN | B-spline | |||
| Modified EfficientKAN | B-spline | ||||
| Modified SechKAN | sech | ||||
| SechKAN (Proposed) | sech | ||||
| Fashion-MNIST | EfficientKAN | B-spline | |||
| Modified EfficientKAN | B-spline |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Function | Formula |
|---|---|
| sech | |
| Gaussian Radial Basis Function (Gaussian RBF) | |
| Derivative of Gaussian | |
| Mexican hat wavelet | |
| Morlet wavelet | |
| Cubic B-spline |
| Basis Function | Runtime (ms) |
|---|---|
| Sech | |
| Gaussian RBF | |
| Derivative of Gaussian | |
| Mexican hat wavelet | |
| Morlet wavelet ( ) | |
| Cubic B-spline |
| Basis Function | Runtime (ms) |
|---|---|
| B-spline (EfficientKAN) | |
| Bare sech (SechKAN) | |
| Parametric sech (SechKAN) | |
| Gaussian RBF (FastKAN) | |
| RSWAF (FasterKAN) |
| Model | Function | Used Params | Loss (MSE) | Runtime (s) |
|---|---|---|---|---|
| BSRBF-KAN | 22 | |||
| FastKAN | 22 | |||
| ReLU-KAN | 22 | |||
| SechKAN | 22 | |||
| EfficientKAN | 22 | |||
| BSRBF-KAN | 22 |
| Model | Val. Rel. L2 | Test Rel. L2 | Runtime (s) | Used Params |
|---|---|---|---|---|
| SechKAN1 | 0.14179 0.01390 | 0.31438 0.02998 | 290.13 2.10 | 38267 |
| SechKAN2 | 0.22064 0.21106 | 0.39155 0.20445 | 298.70 9.58 | 39035 |
| SechKAN3 | 0.20465 0.04908 | 0.33635 0.04548 | 316.47 67.48 | 38267 |
| SechKAN4 | 0.15595 0.03044 | 0.31355 0.03014 | 229.53 50.08 | 38243 |
| SechKAN5 ( Table 7 ) | 0.13664 0.01529 | 0.27695 0.00954 | 266.04 1.87 | 39011 |
| SechKAN6 | 0.33738 0.18304 | 0.41670 0.13447 | 197.55 64.63 | 38243 |
| Model | Val. Rel. L2 | Test Rel. L2 | Runtime (s) | Used Params |
|---|---|---|---|---|
| SechKAN1 ( Table 8 ) | 0.12637 0.00567 | 0.29879 0.01425 | 196.59 18.88 | 39228 |
| SechKAN2 | 0.21322 0.10831 | 0.34455 0.09023 | 209.61 40.14 | 39204 |
| SechKAN3 | 0.29362 0.12529 | 0.39399 0.10176 | 87.34 19.05 | 38436 |
| SechKAN4 | 0.18429 0.11008 | 0.33093 0.10245 | 167.14 66.28 | 38436 |
| Average | 0.20438 | 0.34207 | 165.17 | 38826 |
| Dataset | Network | Acc / Time (s) | Acc / Param |
|---|---|---|---|
| MNIST | MLP | ||
| SechKAN | |||
| SechKAN-CNN | |||
| CNN | |||
| Fashion-MNIST | MLP | ||
| SechKAN |
| Dataset | Model | Network structure | Used Params | MFLOPs | Peak Mem. (MB) |
|---|---|---|---|---|---|
| MNIST Fashion-MNIST | SechKAN | (784, 64, 10) | 52,608 | 0.096 | 23.84 |
| BSRBF-KAN | (784,7,10) | 51,604 | 0.324 | 41.55 | |
| FastKAN | (784,7,10) | 51,621 | 0.108 | 26.59 | |
| FasterKAN | (784,8,10) | 52,400 | 0.086 | 26.64 | |
| EfficientKAN | (784,7,10) | 50,022 | 0.273 | 24.36 | |
| CIFAR-10 | SechKAN | (3072, 64, 10) | 203,632 | 0.501 | 31.84 |
| Dataset | Network | Train. Acc. | Val. Acc. | Val. F1 | Runtime (s) |
|---|---|---|---|---|---|
| MNIST | BSRBF-KAN | ||||
| EfficientKAN | |||||
| FastKAN | |||||
| FasterKAN | |||||
| MLP | |||||
| SechKAN |
| Model | Norm1 | Norm2 | Width | Best Loss | Final Grad. | Max Grad. | Grad. Std. |
|---|---|---|---|---|---|---|---|
| Model1 | None | None | No | ||||
| Model2 | None | None | Yes | ||||
| Model3 | LayerNorm | None | No | ||||
| Model4 | BatchNorm | None | No | ||||
| Model5 | RMSNorm | None | No | ||||
| Model6 | Min–Max | None | No |