KO: Kinetics-inspired Neural Optimizer with PDE Simulation Approaches
Organizations: Shanghai Jiao Tong University
Abstract
The design of effective optimization algorithms for neural networks remains a fundamental challenge, and most existing methods rely on heuristic extensions of gradient-based updates. We introduce KO (Kinetics-inspired Optimizer), a plug-and-play optimization module grounded in kinetic theory and partial differential equations. KO models parameter dynamics as a particle system, augmenting standard gradient updates with stochastic interactions induced by a discretization of the Boltzmann transport equation. This mechanism naturally promotes parameter diversity and mitigates weight condensation, the tendency of parameters to collapse into low-dimensional subspaces, a phenomenon closely associated with degraded generalization. We provide both a rigorous theoretical analysis and a physical interpretation, showing that KO provably increases parameter diversity while preserving convergence guarantees. Extensive experiments on image classification benchmarks (CIFAR-10/100, ImageNet) and large-scale language model pretraining demonstrate that KO consistently improves accuracy over competitive baselines with negligible additional computational cost.
Figures & tables
| Model | CIFAR10 | CIFAR100 |
| ResNet18+SGD | ||
| ResNet18+S.C. | ||
| ResNet18+H.C. | ||
| ResNet18+WCreg | ||
| ResNet34+SGD | ||
| ResNet34+S.C. |
| Model Acc | PIQA | ARC_C | ARC_E | HellaSwag | WinoGrande | OBQA | Avg |
| Qwen2-0.5B+AdamW | 0.575 | 0.227 | 0.371 | 0.273 | 0.498 | 0.246 | 0.365 |
| Qwen2-0.5B+S.C. | 0.582 | 0.241 | 0.385 | 0.276 | 0.522 | 0.268 | 0.379 |
| Qwen2-1.5B+AdamW | 0.602 | 0.236 | 0.439 | 0.310 | 0.495 | 0.256 | 0.390 |
| Qwen2-1.5B+S.C. | 0.615 | 0.258 | 0.450 | 0.313 | 0.535 | 0.266 | 0.406 |
| Model | T-1 ACC.(%) | T-5 ACC.(%) |
| ResNet50+Lamb | 79.70% | 94.53% |
| ResNet50+S.C. | 80.23% | 94.79% |
| ResNet50+H.C. | 79.98% | 94.61% |
| ConvNext_Tiny+AdamW | 82.10% | 96.03% |
| ConvNext_Tiny+S.C. | 82.34% | 96.07% |
| ConvNext_Tiny+H.C. | 82.17% | 96.03% |
| Pretrained Model | IMDB | Snips |
| bert-base-cased+AdamW | 92.52% | 98.14% |
| bert-base-cased+S.C. | 93.80% | 98.43% |
| bert-base-cased+H.C. | 93.42% | 98.43% |
| bert-base-uncased+AdamW | 93.59% | 97.71% |
| bert-base-uncased+S.C. | 94.20% | 97.89% |
| bert-base-uncased+H.C. | 93.76% | 97.73% |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Weight Decay | Init | Mid | End |
| 0 | 0.3867 | 0.9672 | 0.9999 |
| 0.3867 | 0.9676 | 0.9999 | |
| 0.3867 | 0.9245 | 0.9992 |
| Treatment | Task-macro | Sample Accuracy | Correct |
| Historical Standard LoRA | 0.27653 | 27.87% | 1,117 / 4,008 |
| LoRA + KO | 0.28954 | 28.49% | 1,142 / 4,008 |
| KO - baseline | +0.01301 | +0.62% | +25 |
| Treatment | Task-macro | Sample Accuracy | Correct |
| Historical Standard LoRA | 0.36623 | 35.48% | 1,422 / 4,008 |
| LoRA + KO | 0.37491 | 36.40% | 1,459 / 4,008 |
| KO - baseline | +0.00867 | +0.92% | +37 |
| FGSM ( Goodfellow et al., 2015 ) | PGD ( Madry et al., 2019 ) | |
| ResNet18 | 40.46% | 0.08% |
| ResNet18+S.C. | 43.61% | 0.19% |
| ResNet18+H.C. | 46.63% | 0.14% |
| ResNet34 | 42.13% | 0.04% |
| ResNet34+S.C. | 44.72% | 0.23% |
| ResNet34+H.C. | 50.08% | 0.12% |
| 0.0 (Baseline) | 0.1 | 0.2 | 0.3 | 0.4 | 0.5 | 0.6 | 0.7 | 0.8 | 0.9 | |
| Accuracy | 77.38% | 78.23% | 78.80% | 78.30% | 77.98% | 78.92% | 77.52% | 77.87% | 78.01% | 78.16% |
| Gain | - | +0.85% | +1.42% | +0.92% | +0.60% | +1.54% | +0.14% | +0.49% | +0.63% | +0.78% |
| Vanilla | Vanilla + WD | S.C. | S.C. + WD | H.C. | H.C. + WD | |
| ResNet18 | 95.02% 0.29% | 95.07% 0.18% | 95.74% 0.09% | 95.74% 0.05% | 95.42% 0.13% | 95.58% 0.07% |
| ResNet34 | 95.11% 0.32% | 95.14% 0.37% | 95.52% 0.12% | 95.76% 0.08% | 95.33% 0.18% | 95.56% 0.14% |
| ResNet50 | 95.31% 0.27% | 95.37% 0.38% | 95.62% 0.08% | 95.83% 0.16% | 95.42% 0.11% | 95.57% 0.05% |
| Concept | Kinetic Theory (Physical System) | Neural Network Optimization |
| Agent | A single particle | A single neuron (weight vector ) |
| Position | Spatial coordinates of the particle | Weight vector |
| Velocity | Particle velocity vector | (Negative) gradient (update direction) |
| Interaction | Stochastic collisions | Gradient modification (hard/soft collision terms) |
| System Dynamics | Boltzmann Transport Equation | Optimization trajectory (gradient descent + collision) |