Explicit residual connections of the form (x+f(x)), often combined with normalization layers, have become a standard strategy for training very deep neural networks. However, residual addition primarily provides an algebraic shortcut for gradient propagation, while leaving the evolution of feature geometry across layers largely unconstrained. We introduce Learnable Lens Networks (LLN), a physics-inspired architecture that replaces direct feature-space residual accumulation with learnable optical transport in an augmented position-angle phase space. Each layer alternates between free propagation, which provides an implicit transport path, and a learnable lens field that performs nonlinear trajectory transformation and focusing. Theoretically, we establish that LLN transport is globally invertible and volume-preserving for any differentiable lens field, with the implemented coordinate-wise Gaussian transport further satisfying symplecticity. Importantly, these structural constraints do not limit expressivity: with unrestricted embeddings and readouts, LLN retain universal approximation of continuous end-to-end maps. Experiments across diverse dynamical systems demonstrate that LLN improves long-horizon prediction while using substantially fewer parameters than same-depth comparators. Further analysis reveals stable depth-wise gradient transport and interpretable learned dynamics under the coupled propagation and refraction design.
Figures & tables
Figure 1: LLN architecture overview. (a) A conventional residual block adds a nonlinear transform in the same feature space, which can accumulate forward variance when normalization is removed. (b) An LLN layer separates feature position y and angle θ ; free propagation provides an implicit physical residual path, and a learnable lens field refracts the angle. (c) Depth repeats the propagate-refract optical map, yielding a phase-space layer Jacobian with det(Jl)=1 . Dashed connectors indicate how the single-layer operations in (b) are repeated across the stack in (c).
Figure 2: All-model rollout error across network depths. Curves show the mean normalized rollout MSE over ten runs for LLN and five non-LLN architectures at each evaluated depth. MLP denotes the plain multilayer perceptron, Res denotes the residual MLP, Res-S denotes the scaled residual MLP, Res-BN denotes the residual MLP with batch normalization (BN), and LS denotes LayerScale. Lower is better.
System
Family
LLN MSE
Selected non-LLN
Comparator MSE
MSE ratio
Param ratio
FLOP ratio
Lorenz-24
chaotic flow
6.170±0.953
Plain MLP
7.516±0.010
0.821
28.5%
29.2%
Lorenz-28
chaotic flow
0.920±0.545
Scaled ResMLP
0.831±0.942
1.106
6.4%
29.0%
SpringChain-32
mechanical
1.070±0.379
Plain MLP
1.193±0.001
0.897
38.9%
32.0%
SpringChain-48
mechanical
1.087±0.017
Plain MLP
1.135±0.001
0.958
57.5%
33.4%
Burgers-32
fluid
0.537±0.155
Plain MLP
2.096±0.020
0.256
5.5%
30.6%
Burgers-48
fluid
0.550±0.122
LayerScale
1.973±0.030
0.279
5.9%
31.1%
Table 1: Depth-40 selected-deployment rollout performance. Values denote normalized MSE (mean ± std over 10 seeds) over a 500-step free-running rollout. The comparator is the lowest-MSE member of the designated non-LLN suite for each protocol block; all paired entries match depth, trajectory split, seed set, and evaluation metric. The MSE and parameter ratios use the displayed deployments, whereas the floating-point operation (FLOP) ratio is a strict matched-width audit with both models set to D=40 and h=128 . Lower is better; bold marks the lower paired rollout MSE.
Figure 3: Matched layerwise gradient transport at D=40 and h=128 . Curves show the mean log10(gl/g40) over 10 initialization seeds on five representative systems; bands show one standard deviation. The lower panels magnify Scaled ResMLP and LLN near neutral transport.
System
Plain MLP
ResMLP
Scaled ResMLP
LLN
Lorenz-24
−15.298±0.222
3.140±0.152
0.087±0.048
0.000±0.008
Lorenz-28
−15.298±0.219
3.141±0.151
0.088±0.048
0.001±0.008
SpringChain-32
−15.270±0.257
3.741±0.064
0.286±0.034
0.003±0.005
SpringChain-48
−15.306±0.214
3.818±0.079
0.333±0.034
0.003±0.008
Burgers-32
−15.326±0.265
3.620±0.085
0.217±0.032
0.004±0.006
Burgers-48
−15.355±0.207
3.697±0.093
0.251±0.043
0.004±0.007
Table 2: Matched gradient transport at initialization. Values are log10(g1/g40) (mean ± s.d. over 10 seeds). Zero is ideal; bold marks the closest value to zero.
Figure 4: Low-dimensional optical case study. (a) A deterministic boundary error and two nearest correct references. (b) LLN position paths. (c) Plain MLP paths for the same inputs.
Figure 5: Layerwise LLN refraction and readout. Paths trace the class-1 reference, boundary error, and class-0 reference through the wrong sample’s largest-turn layers (L1, L4, L7); the right table lists all-layer turns.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Evidence block
Systems
Selection/configuration
Protocol
Table 1 : main rollout
L24/28; SC32/48; B32/48; AD32/48; RD32/48
Selected LLN configuration versus the lowest-MSE designated same-depth non-LLN comparator; LLN configurations are listed in Table 4 .
R0, SC-C, L28-C
Table B.2 : multi-trajectory confirmation
L24/28; SC32/48; B32/48; AD32/48; RD32/48
Main-table LLN and comparator configurations fixed before confirmatory retraining; reported estimates use training seeds 11-16.
MT
Figure B.3 : rollout curves
RD48, AD48, B48, L24
Configurations selected using seeds 0-2 and evaluated using disjoint seeds 3-9; D=40 , K=12 , s=1 , and learned θ0 . This is a representative visualization rather than an all-system claim.
C0
Table 2 : gradient transport
L24/28; SC32/48; B32/48; AD32/48; RD32/48
All models use D=40 and h=128 ; LLN additionally uses T=1 , K=12 , s=1 , and learned θ0 .
G0
Table 6 : structure ablation
L24/28; SC32/48; B32/48; AD32/48; RD32/48
Main-table LLN configurations; the three ablations set s=0 ( w/o lens ), θ0=0 , or L=0 ( w/o propagation ), respectively.
R0
Table 7 : all depths
L24/28; SC32/48; B32/48; AD32/48; RD32/48
D∈{5,10,20,40} ; baselines use h=128 . The LLN grid varies T and h with default K=12 and s=1 ; completion rows use their recorded system-specific configurations.
R0, SC-C, L28-C
Appendix
Table 3: Protocol index for the empirical evidence. System abbreviations are L=Lorenz, SC=SpringChain, B=Burgers, AD=AdvecDiff, and RD=ReactionDiff.
System
T
h
K
s
Lorenz-24
3
384
12
1
Lorenz-28
0.75
128
8
0.075
SpringChain-32
1
512
8
1.25
SpringChain-48
0.25
512
12
1
Burgers-32
2
64
12
1
Burgers-48
2
64
12
1
Appendix
Table 4: Selected LLN deployments for the main D=40 rollout results. T denotes the optical horizon, h the hidden width, K the number of Gaussian basis functions, and s the fixed lens-output scale. All configurations use a learned initial angle θ0 .
System
LLN MSE
Comparator MSE
MSE ratio [95% CI]
Lorenz-24
2.027
0.554
3.656[1.180,13.444]
Lorenz-28
2.760
0.448
6.165[4.769,7.947]
SpringChain-32
1.919
0.935
2.051[1.560,2.741]
SpringChain-48
0.053
0.950
0.056[0.037,0.083]
Burgers-32
0.289
0.999
0.289[0.177,0.482]
Burgers-48
0.199
4.922
0.040[0.00109,0.451]
Appendix
Table 5: Frozen multi-trajectory confirmation across all ten D=40 systems. MSE values are geometric means over six confirmatory training seeds, four unseen initial-condition trajectories, and two non-overlapping 500-step windows per trajectory. Ratio intervals are 95% seed-cluster bootstrap intervals. Lower is better; bold marks the lower paired MSE, and bold ratios mark intervals entirely below one.
Figure 6: Temporal rollout error at depth D=40 . Lines show geometric-mean normalized state MSE over seven confirmation seeds; bands are 95% seed-bootstrap intervals. Lower is better.
System
Full LLN
Without lens refraction
Zero initial angle
Without free propagation
Lorenz-24
6.18±0.98
8.54±3.94
9.95±3.43
∞
Lorenz-28
1.03±0.48
1.37±0.66
5.6×1010
1.0×105
SpringChain-32
1.22±0.18
1.22±0.13
1.19±0.01
1.17±0.01
SpringChain-48
1.09±0.02
1.10±0.01
1.11±0.01
1.11±0.00
Burgers-32
0.54±0.16
0.79±0.26
1.78±1.00
82.25±253.94
Burgers-48
0.55±0.12
0.79±0.27
1.22±0.55
1.9×1016
Appendix
Table 6: Complete LLN structure ablation across all ten systems at D=40 . Values are 500-step normalized rollout MSE (mean ± s.d. over 10 independent stochastic runs). Extremely large divergent values are reported in scientific notation. Within each system, all variants use the same width, optical horizon, Gaussian basis count, optimizer, and data protocol; each ablation modifies only the indicated architectural component. Lower is better, and bold marks the best value in each row.
D
System
LLN
MLP
Res
Res-S
Res-BN
LS
MSE r.
Param r.
FLOP r.
5
Lorenz-24
6.778±2.160
7.510 ± 0.066
∞
7.348 ± 0.805
7.813 ± 1.286
7.562 ± 0.294
0.922
0.054
0.301
5
Lorenz-28
1.692 ± 1.300
0.292±0.165
1.742 ± 0.907
0.637 ± 0.862
1.313 ± 1.089
1.236 ± 0.332
5.790
0.078
0.303
5
SpringChain-32
1.157±0.033
1.194 ± 0.003
1.187 ± 0.018
1.191 ± 0.009
1.461 ± 0.065
1.193 ± 0.006
0.975
0.985
0.490
5
SpringChain-48
1.151±0.112
∞
∞
∞
1.448 ± 0.201
∞
0.794
0.826
0.560
5
Burgers-32
0.969±0.319
2.090 ± 0.024
2.250 ± 0.127
2.138 ± 0.048
2.244 ± 0.270
2.105 ± 0.037
0.464
0.112
0.401
5
Burgers-48
0.803±0.279
1.984 ± 0.016
2.185 ± 0.170
2.026 ± 0.051
2.218 ± 0.251
1.977 ± 0.030
0.406
0.138
0.445
Appendix
Table 7: All-depth rollout comparison on the ten-system benchmark. Values are normalized rollout MSE (mean ± s.d. over 10 seeds); the final columns report LLN-to-comparator MSE, parameter, and matched arithmetic-FLOP ratios. LS denotes LayerScale and “r.” denotes ratio. Lower is better; bold marks the lowest finite rollout MSE among the compared models in each row.
Rank
Test index
Error
Confidence
Max. turn ( ∘ )
Peak layer
0
183
1→0
0.547
16.62
L1
1
94
0→1
0.560
18.52
L1
2
260
1→0
0.592
17.38
L1
3
199
1→0
0.624
13.93
L1
4
107
0→1
0.661
16.11
L7
5
366
0→1
0.726
17.31
L7
Appendix
Table 8: Complete index of LLN errors in the fixed two-moons case study. Errors are written as true class → predicted class. Rank orders the errors by increasing predicted-class confidence, and the peak layer maximizes Δαl .
Figure 7: Large overview of all seven LLN error paths and turn statistics. (a) Input-space locations of the seven deterministic LLN errors. (b) LLN position trajectories in the full-test PCA basis; numbers denote confidence rank and stars mark each error path’s peak-turn layer. (c) Layerwise wrong-sample turn angles for all error ranks. (d) Full-test maximum-turn distribution comparing LLN errors with correct samples.
Shanghai Institute for Mathematics and Interdisciplinary Sciences (SIMIS) Shanghai 200433, China · Academy of Mathematics and Systems Science Chinese Academy of Sciences Beijing 100190, China · School of Mathematical Sciences2026 University of Chinese Academy of Sciences Beijing 100049, China