cs.LGMay 20, 2025

KO: Kinetics-inspired Neural Optimizer with PDE Simulation Approaches

Authors: Mingquan Feng, Yixin Huang, Yifan Fu, Shaobo Wang, Junchi Yan

Organizations: Shanghai Jiao Tong University

Abstract

The design of effective optimization algorithms for neural networks remains a fundamental challenge, and most existing methods rely on heuristic extensions of gradient-based updates. We introduce KO (Kinetics-inspired Optimizer), a plug-and-play optimization module grounded in kinetic theory and partial differential equations. KO models parameter dynamics as a particle system, augmenting standard gradient updates with stochastic interactions induced by a discretization of the Boltzmann transport equation. This mechanism naturally promotes parameter diversity and mitigates weight condensation, the tendency of parameters to collapse into low-dimensional subspaces, a phenomenon closely associated with degraded generalization. We provide both a rigorous theoretical analysis and a physical interpretation, showing that KO provably increases parameter diversity while preserving convergence guarantees. Extensive experiments on image classification benchmarks (CIFAR-10/100, ImageNet) and large-scale language model pretraining demonstrate that KO consistently improves accuracy over competitive baselines with negligible additional computational cost.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Anatomy of Implicit Bias: Information Allocation in Neural Network Training

    Jul 8, 2026Zhang Gongyue, Wang Zhiyong, Liu Donghan +3Neural Network OptimizationOptimizer Design

  2. Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization

    Sep 3, 2025Wu Lin, Scott C. Lowe, Felix Dangel +3Neural Network OptimizationKullback-Leibler Divergence