Explicit residual connections of the form (x+f(x)), often combined with normalization layers, have become a standard strategy for training very deep neural networks. However, residual addition primarily provides an algebraic shortcut for gradient propagation, while leaving the evolution of feature geometry across layers largely unconstrained. We introduce Learnable Lens Networks (LLN), a physics-inspired architecture that replaces direct feature-space residual accumulation with learnable optical transport in an augmented position-angle phase space. Each layer alternates between free propagation, which provides an implicit transport path, and a learnable lens field that performs nonlinear trajectory transformation and focusing. Theoretically, we establish that LLN transport is globally invertible and volume-preserving for any differentiable lens field, with the implemented coordinate-wise Gaussian transport further satisfying symplecticity. Importantly, these structural constraints do not limit expressivity: with unrestricted embeddings and readouts, LLN retain universal approximation of continuous end-to-end maps. Experiments across diverse dynamical systems demonstrate that LLN improves long-horizon prediction while using substantially fewer parameters than same-depth comparators. Further analysis reveals stable depth-wise gradient transport and interpretable learned dynamics under the coupled propagation and refraction design.
Figures & tables
Figure 1: LLN architecture overview. (a) A conventional residual block adds a nonlinear transform in the same feature space, which can accumulate forward variance when normalization is removed. (b) An LLN layer separates feature position y and angle θ ; free propagation provides an implicit physical residual path, and a learnable lens field refracts the angle. (c) Depth repeats the propagate-refract optical map, yielding a phase-space layer Jacobian with det(Jl)=1 . Dashed connectors indicate how the single-layer operations in (b) are repeated across the stack in (c).
Figure 2: All-model rollout error across network depths. Curves show the mean normalized rollout MSE over ten runs for LLN and five non-LLN architectures at each evaluated depth. MLP denotes the plain multilayer perceptron, Res denotes the residual MLP, Res-S denotes the scaled residual MLP, Res-BN denotes the residual MLP with batch normalization (BN), and LS denotes LayerScale. Lower is better.
System
Family
LLN MSE
Selected non-LLN
Comparator MSE
MSE ratio
Param ratio
FLOP ratio
Lorenz-24
chaotic flow
6.170±0.953
Plain MLP
7.516±0.010
0.821
28.5%
29.2%
Lorenz-28
chaotic flow
0.920±0.545
Scaled ResMLP
0.831±0.942
1.106
6.4%
29.0%
SpringChain-32
mechanical
1.070±0.379
Plain MLP
1.193±0.001
0.897
38.9%
32.0%
SpringChain-48
mechanical
1.087±0.017
Plain MLP
1.135±0.001
0.958
57.5%
33.4%
Burgers-32
fluid
0.537±0.155
Plain MLP
2.096±0.020
0.256
5.5%
30.6%
Burgers-48
fluid
0.550±0.122
LayerScale
1.973±0.030
0.279
5.9%
31.1%
Table 1: Depth-40 selected-deployment rollout performance. Values denote normalized MSE (mean ± std over 10 seeds) over a 500-step free-running rollout. The comparator is the lowest-MSE member of the designated non-LLN suite for each protocol block; all paired entries match depth, trajectory split, seed set, and evaluation metric. The MSE and parameter ratios use the displayed deployments, whereas the floating-point operation (FLOP) ratio is a strict matched-width audit with both models set to D=40 and h=128 . Lower is better; bold marks the lower paired rollout MSE.
Figure 3: Matched layerwise gradient transport at D=40 and h=128 . Curves show the mean log10(gl/g40) over 10 initialization seeds on five representative systems; bands show one standard deviation. The lower panels magnify Scaled ResMLP and LLN near neutral transport.
System
Plain MLP
ResMLP
Scaled ResMLP
LLN
Lorenz-24
−15.298±0.222
3.140±0.152
0.087±0.048
0.000±0.008
Lorenz-28
−15.298±0.219
3.141±0.151
0.088±0.048
0.001±0.008
SpringChain-32
−15.270±0.257
3.741±0.064
0.286±0.034
0.003±0.005
SpringChain-48
−15.306±0.214
3.818±0.079
0.333±0.034
0.003±0.008
Burgers-32
−15.326±0.265
3.620±0.085
0.217±0.032
0.004±0.006
Burgers-48
−15.355±0.207
3.697±0.093
0.251±0.043
0.004±0.007
Table 2: Matched gradient transport at initialization. Values are log10(g1/g40) (mean ± s.d. over 10 seeds). Zero is ideal; bold marks the closest value to zero.
Figure 4: Low-dimensional optical case study. (a) A deterministic boundary error and two nearest correct references. (b) LLN position paths. (c) Plain MLP paths for the same inputs.
Figure 5: Layerwise LLN refraction and readout. Paths trace the class-1 reference, boundary error, and class-0 reference through the wrong sample’s largest-turn layers (L1, L4, L7); the right table lists all-layer turns.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Evidence block
Systems
Selection/configuration
Protocol
Table 1 : main rollout
L24/28; SC32/48; B32/48; AD32/48; RD32/48
Selected LLN configuration versus the lowest-MSE designated same-depth non-LLN comparator; LLN configurations are listed in Table 4 .
R0, SC-C, L28-C
Table B.2 : multi-trajectory confirmation
L24/28; SC32/48; B32/48; AD32/48; RD32/48
Main-table LLN and comparator configurations fixed before confirmatory retraining; reported estimates use training seeds 11-16.
MT
Figure B.3 : rollout curves
RD48, AD48, B48, L24
Configurations selected using seeds 0-2 and evaluated using disjoint seeds 3-9; D=40 , K=12 , s=1 , and learned θ0 . This is a representative visualization rather than an all-system claim.
C0
Table 2 : gradient transport
L24/28; SC32/48; B32/48; AD32/48; RD32/48
All models use D=40 and h=128 ; LLN additionally uses T=1 , K=12 , s=1 , and learned θ0 .
G0
Table 6 : structure ablation
L24/28; SC32/48; B32/48; AD32/48; RD32/48
Main-table LLN configurations; the three ablations set s=0 ( w/o lens ), θ0=0 , or L=0 ( w/o propagation ), respectively.
R0
Table 7 : all depths
L24/28; SC32/48; B32/48; AD32/48; RD32/48
D∈{5,10,20,40} ; baselines use h=128 . The LLN grid varies T and h with default K=12 and s=1 ; completion rows use their recorded system-specific configurations.
R0, SC-C, L28-C
Appendix
Table 3: Protocol index for the empirical evidence. System abbreviations are L=Lorenz, SC=SpringChain, B=Burgers, AD=AdvecDiff, and RD=ReactionDiff.
System
T
h
K
s
Lorenz-24
3
384
12
1
Lorenz-28
0.75
128
8
0.075
SpringChain-32
1
512
8
1.25
SpringChain-48
0.25
512
12
1
Burgers-32
2
64
12
1
Burgers-48
2
64
12
1
Appendix
Table 4: Selected LLN deployments for the main D=40 rollout results. T denotes the optical horizon, h the hidden width, K the number of Gaussian basis functions, and s the fixed lens-output scale. All configurations use a learned initial angle θ0 .
System
LLN MSE
Comparator MSE
MSE ratio [95% CI]
Lorenz-24
2.027
0.554
3.656[1.180,13.444]
Lorenz-28
2.760
0.448
6.165[4.769,7.947]
SpringChain-32
1.919
0.935
2.051[1.560,2.741]
SpringChain-48
0.053
0.950
0.056[0.037,0.083]
Burgers-32
0.289
0.999
0.289[0.177,0.482]
Burgers-48
0.199
4.922
0.040[0.00109,0.451]
Appendix
Table 5: Frozen multi-trajectory confirmation across all ten D=40 systems. MSE values are geometric means over six confirmatory training seeds, four unseen initial-condition trajectories, and two non-overlapping 500-step windows per trajectory. Ratio intervals are 95% seed-cluster bootstrap intervals. Lower is better; bold marks the lower paired MSE, and bold ratios mark intervals entirely below one.
Figure 6: Temporal rollout error at depth D=40 . Lines show geometric-mean normalized state MSE over seven confirmation seeds; bands are 95% seed-bootstrap intervals. Lower is better.
System
Full LLN
Without lens refraction
Zero initial angle
Without free propagation
Lorenz-24
6.18±0.98
8.54±3.94
9.95±3.43
∞
Lorenz-28
1.03±0.48
1.37±0.66
5.6×1010
1.0×105
SpringChain-32
1.22±0.18
1.22±0.13
1.19±0.01
1.17±0.01
SpringChain-48
1.09±0.02
1.10±0.01
1.11±0.01
1.11±0.00
Burgers-32
0.54±0.16
0.79±0.26
1.78±1.00
82.25±253.94
Burgers-48
0.55±0.12
0.79±0.27
1.22±0.55
1.9×1016
Appendix
Table 6: Complete LLN structure ablation across all ten systems at D=40 . Values are 500-step normalized rollout MSE (mean ± s.d. over 10 independent stochastic runs). Extremely large divergent values are reported in scientific notation. Within each system, all variants use the same width, optical horizon, Gaussian basis count, optimizer, and data protocol; each ablation modifies only the indicated architectural component. Lower is better, and bold marks the best value in each row.
D
System
LLN
MLP
Res
Res-S
Res-BN
LS
MSE r.
Param r.
FLOP r.
5
Lorenz-24
6.778±2.160
7.510 ± 0.066
∞
7.348 ± 0.805
7.813 ± 1.286
7.562 ± 0.294
0.922
0.054
0.301
5
Lorenz-28
1.692 ± 1.300
0.292±0.165
1.742 ± 0.907
0.637 ± 0.862
1.313 ± 1.089
1.236 ± 0.332
5.790
0.078
0.303
5
SpringChain-32
1.157±0.033
1.194 ± 0.003
1.187 ± 0.018
1.191 ± 0.009
1.461 ± 0.065
1.193 ± 0.006
0.975
0.985
0.490
5
SpringChain-48
1.151±0.112
∞
∞
∞
1.448 ± 0.201
∞
0.794
0.826
0.560
5
Burgers-32
0.969±0.319
2.090 ± 0.024
2.250 ± 0.127
2.138 ± 0.048
2.244 ± 0.270
2.105 ± 0.037
0.464
0.112
0.401
5
Burgers-48
0.803±0.279
1.984 ± 0.016
2.185 ± 0.170
2.026 ± 0.051
2.218 ± 0.251
1.977 ± 0.030
0.406
0.138
0.445
Appendix
Table 7: All-depth rollout comparison on the ten-system benchmark. Values are normalized rollout MSE (mean ± s.d. over 10 seeds); the final columns report LLN-to-comparator MSE, parameter, and matched arithmetic-FLOP ratios. LS denotes LayerScale and “r.” denotes ratio. Lower is better; bold marks the lowest finite rollout MSE among the compared models in each row.
Rank
Test index
Error
Confidence
Max. turn ( ∘ )
Peak layer
0
183
1→0
0.547
16.62
L1
1
94
0→1
0.560
18.52
L1
2
260
1→0
0.592
17.38
L1
3
199
1→0
0.624
13.93
L1
4
107
0→1
0.661
16.11
L7
5
366
0→1
0.726
17.31
L7
Appendix
Table 8: Complete index of LLN errors in the fixed two-moons case study. Errors are written as true class → predicted class. Rank orders the errors by increasing predicted-class confidence, and the peak layer maximizes Δαl .
Figure 7: Large overview of all seven LLN error paths and turn statistics. (a) Input-space locations of the seven deterministic LLN errors. (b) LLN position trajectories in the full-test PCA basis; numbers denote confidence rank and stars mark each error path’s peak-turn layer. (c) Layerwise wrong-sample turn angles for all error ranks. (d) Full-test maximum-turn distribution comparing LLN errors with correct samples.
Forecasting the long-horizon evolution of mechanical systems from position-only observations is a pivotal yet difficult task, as hidden velocities and trajectory-specific physical properties must be inferred simultaneously. Although physics-guided neural networks like Lagrangian Neural Networks (LNNs) guarantee physical plausibility, they generally require complete state inputs and lack adaptability to changing system parameters. To break these limitations, we introduce History-informed Lagrangian Neural Networks (HiLNN). Grounded in the insight that temporal position sequences implicitly encode underlying dynamics, HiLNN employs a recurrent encoder to extract a latent context from history. This context not only reconstructs the unobserved initial velocity but also adaptively modulates the mass matrix, potential energy, and damping coefficients of a structured Lagrangian system. By leveraging a differentiable RK4 rollout scheme, the entire pipeline is optimized end-to-end under multi-step trajectory supervision and energy-consistency regularization. Empirical evaluations across conservative, dissipative, and heterogeneous variable-parameter systems show that HiLNN delivers superior long-term prediction accuracy and maintains precise energy profiles compared to state-of-the-art baselines. The source code is publicly available at https://github.com/yingtian22/History-informed-LNN.
Tianshuo Zhang, Xianglei Xing, Wenzhe Zhai +2
College of Intelligent Systems Science and Engineering, Harbin Engineering University, Harbin 150001, China
The Universal Approximation Theorem (UAT) guarantees universal function approximation but does not explain how residual models distribute approximation across layers. We reframe residual networks as a layer-wise approximation process that builds an approximation trajectory from input to target, and prove the existence of progressive trajectories where error decreases monotonically with depth. It reveals that residual networks can implement structured, step-by-step refinement rather than end-to-end (E2E) black-box mapping. Building on this, we propose Layer-wise Progressive Approximation (LPA), a theoretically grounded training principle that explicitly aligns each layer with its residual target to realize such trajectories. LPA is architecture-agnostic: we observe progressive behavior in residual FNNs, ResNets, and Transformers across tasks including complex surface fitting, image classification, and NLP with LLMs for generation and classification. Crucially, this enables ``train once, use N models": a single network yields useful predictions at every depth, supporting efficient shallow inference without retraining. Our work unifies approximation theory with practical deep learning, providing a new lens on representation learning and a flexible framework for multi-depth deployment. The source code will be released unpon acceptance at https://(open_upon_acceptance).
Wei Wang, Xiao-Yong Wei, Qing Li
Department of Computing, the Hong Kong Polytechnic University.
Recent studies revealed the mathematical connection between deep neural networks (DNNs) and dynamic systems. However, the specific dynamics that DNNs, especially deep residual networks (ResNets), tend to learn during training remain insufficiently characterized. To this end, we model the forward propagation of deep residual networks using continuity equations, in which the measure is conserved and infinite curves in the measure space connect the input distribution to the output one of a ResNet. We find ResNets with L2 regularization attempt to learn the geodesic curve in the Wasserstein space, induced by the optimal transport map. Compared with plain networks, ResNets can better approximate the geodesic curve, which explains why ResNets can be optimized and generalize better. Numerical experiments show that the data tracks of a ResNet tend to be line-shaped in terms of the line-shape score, and the map learned by a ResNet is closer to the optimal transport map in terms of the optimal transport score. In a word, we conclude that ResNets learn the geodesic curve in the Wasserstein space and discretely engineer the data transformation in high-dimensional spaces.
Kuo Gai, Shihua Zhang
Shanghai Institute for Mathematics and Interdisciplinary Sciences (SIMIS) Shanghai 200433, China · Academy of Mathematics and Systems Science Chinese Academy of Sciences Beijing 100190, China · School of Mathematical Sciences2026 University of Chinese Academy of Sciences Beijing 100049, China