Shallow models are special cases of deep models, and deep models theoretically have the potential to outperform the shallow ones. However, the existing empirical asset pricing literature provides strong benchmarks for shallow models. Residual learning allows neural network models in asset pricing to go deeper by preserving and refining their shallow counterparts. The out-of-sample Sharpe ratio for value-weighted long-short portfolios of deep residual models (2.07) is higher than that for the corresponding shallow ones (1.92) and more than twice that of the deep feedforward models (0.89). We show that model depth is a source of additional economic value in asset pricing. Residual learning can be used to deepen other neural-network-based asset pricing models if they contain intermediate layers. Our design also provides one way to scale asset pricing models, making native "large asset pricing models" more feasible.
Figures & tables
Figure 1: From Feedforward Networks to Residual Neural Networks. The left panel shows a feedforward block, where FNN denotes a feedforward neural network. The right panel shows a residual block. A feedforward block performs a simple function mapping, whereas a residual block keeps the input x intact through a shortcut path and adds a learned residual update E(x) from the residual branch.
Figure 2: Simplified Building Blocks. Circles represent hidden states; arrows represent sequences of operations; text boxes, e.g., FNN and BatchNorm, represent functions. The drawings show the core building blocks of the four models when the input and output dimensions of the building block are the same. For projections and technical formulations, see Appendix A.2 .
Figure 3: Expanding Training Window. Each row represents the information observable in one forecast origin year. Time flows along the y-axis, from bottom to top, and the x-axis represents the sample years observable at the current forecast origin year. The blue points in each row indicate the data in the training sample for that forecast origin year; the orange point indicates the corresponding out-of-sample test sample. The model parameters are re-estimated for the latest forecast origin year to adapt to the latest market conditions. Forecast origins from January through December are used to forecast returns from February through the following January.
Table 1: Empirical Design
Table 2: Long-Short Portfolio Performance by Depth Group and Model Family
Figure 4: Sharpe Ratio Differences. Each grid entry represents a width and depth specification within one comparison set. Each entry reports the out-of-sample Sharpe ratio difference between the residual model and the feedforward model for value-weighted long-short portfolios. Green entries favor the residual model and red entries favor the feedforward model. Dashed green crosses indicate specifications in which the ResNet forecast remains valid while the NN forecast collapses. The diagonal line on the left shows the minimal depth for each width schedule.
Figure 5: Realized Returns, Volatility, and Sharpe Ratios across Deciles. Solid lines report mean monthly realized returns on the left axis, dashed lines report mean volatility of monthly realized returns on the right axis, and the bars report annualized Sharpe ratios. The statistics are calculated from monthly value-weighted portfolios and then averaged across completed width and depth specifications shared by both model families in each comparison set.
NN
Shallow
Medium
Deep
Rank
Fcst
Mean
SD
SR
Fcst
Mean
SD
SR
Fcst
Mean
SD
SR
Low
-2.74
-0.56
8.54
-0.23
-2.18
-0.53
8.50
-0.21
-1.00
-0.42
8.08
-0.18
2
-1.34
0.07
6.62
0.04
-1.10
0.07
6.68
0.04
-0.44
0.14
6.58
0.08
3
-0.59
0.44
5.68
0.27
-0.49
0.46
5.72
0.28
-0.10
0.45
5.62
0.28
4
-0.07
0.53
5.15
0.36
-0.05
0.53
5.17
0.36
0.16
0.54
5.10
0.37
5
0.35
0.67
4.73
0.49
0.30
0.65
4.77
0.47
0.39
0.66
4.78
0.48
Table 3: Decile Portfolio Performance by Depth Group and Model Family
Table 4: Forecast Preservation and Refinement by Depth Group
Figure 6: Forecast Preservation and Refinement by Depth. Solid lines report the preservation score Pt relative to the corresponding shallow anchor model on the left axis. Dashed lines report the revision accuracy Ct+1 (%) relative to the corresponding shallow anchor model on the right axis. Revision accuracy larger than zero indicates refinement, while accuracy less than zero indicates deterioration. Green and red indicate the residual and feedforward models, respectively, in each panel.
ResNet
NN
Theme
ΔPt
ΔCt+1/%
ΔVt+1/bps
ΔPt
ΔCt+1/%
ΔVt+1/bps
Accruals
−0.001∗∗∗
0.02
0.05
−0.001
0.11
0.01
Debt Issuance
0.000
0.04
0.02
0.001
−0.01
0.11
Investment
−0.001∗
0.01
−0.23
−0.001
−0.16
0.07
Low Leverage
−0.003∗∗∗
0.49∗∗∗
0.40
−0.011∗∗∗
0.05
0.57
Low Risk
0.002∗∗
−0.29
−0.65
−0.001
−0.53
0.54
Table 5: Forecast Preservation and Revision Losses by Theme
Table 6: Market Capitalization Exclusion Tests
Table 7: Transaction Cost Tests
Table 8: NBER Business Cycle Tests
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: ResNet+ Forecasting Blocks. The left panel shows Type I , which adds the learned residual update Ed(hd−1) to the identity shortcut path when the hidden width is unchanged. Type II shows a width change with additional layers: P denotes Qd,m in equation ( A.4 ), and FNN denotes the feedforward branch Ed . The circled “+” denotes elementwise addition.
No.
Acronym
Description
Block
Source
Frequency
1
niq_su
Standardized earnings surprise
Firm chars.
JKP panel
Monthly aligned
2
ret_6_1
Price momentum t-6 to t-1
Firm chars.
JKP panel
Monthly aligned
3
ret_12_1
12-month price momentum
Firm chars.
JKP panel
Monthly aligned
4
saleq_su
Standardized revenue surprise
Firm chars.
JKP panel
Monthly aligned
5
tax_gr1a
Tax expense surprise
Firm chars.
JKP panel
Monthly aligned
6
ni_inc8q
Consecutive earnings increases
Firm chars.
JKP panel
Monthly aligned
Appendix
Table 9: Predictors for One-Month-Ahead Return Forecasts
Model
Pt
Ct+1/%
Vt+1S/bps
Vt+1D/bps
Model
Pt
Ct+1/%
Vt+1S/bps
Vt+1D/bps
Medium
ResNet
0.93
0.32
56.72∗∗∗
55.78∗∗∗
ResNet+
0.81
0.72∗∗∗
62.62∗∗∗
63.31∗∗∗
NN
0.87
−0.74∗
56.90∗∗∗
54.83∗∗∗
NN+
0.75
-0.04
62.62∗∗∗
46.73∗∗∗
Difference
0.06∗∗∗
1.06∗∗
-0.18
0.95
Difference
0.06∗∗∗
0.75
0.00
16.58∗∗∗
Deep
ResNet
0.92
0.58∗∗
56.72∗∗∗
55.58∗∗∗
ResNet+
0.79
1.12∗∗∗
62.62∗∗∗
67.54∗∗∗
NN
0.84
-0.99
57.17∗∗∗
54.53∗∗∗
NN+
0.64
−1.88∗∗∗
62.62∗∗∗
28.17∗∗∗
Difference
0.08∗∗∗
1.58∗∗
-0.46
1.04
Difference
0.15∗∗∗
2.99∗∗∗
0.00
39.36∗∗∗
Appendix
Table 10: Forecast Preservation and Ranking Economic Values by Depth Group
ResNet
NN
Theme
ΔPt
ΔCt+1/%
ΔVt+1S/bps
ΔVt+1D/bps
ΔPt
ΔCt+1/%
ΔVt+1S/bps
ΔVt+1D/bps
Accruals
−0.001∗∗∗
0.02
−0.13
−0.08
−0.001
0.11
−0.14
−0.13
Debt Issuance
0.000
0.04
0.77
0.79
0.001
−0.01
0.71
0.83
Investment
−0.001∗
0.01
1.11
0.88
−0.001
−0.16
1.05
1.12
Low Leverage
−0.003∗∗∗
0.49∗∗∗
1.54
1.94
−0.011∗∗∗
0.05
1.27
1.84
Low Risk
0.002∗∗
−0.29
4.02
3.37
−0.001
−0.53
4.05
4.58
Appendix
Table 11: Shallow and Deep Ranking Economic Value Losses by Theme
Model
Depth
Fallback
VW SR
Fcst SD
Unique fcsts
Min hidden SD
NN s1d3
3
0
0.87
2.1×10−2
77,387
1.4×10−3
NN s1d7
7
443
0.21
0
1
0
ResNet s1d3
3
0
0.88
2.8×10−2
77,417
1.6×10−1
ResNet s1d7
7
0
0.90
2.4×10−2
77,440
2.0×10−1
Appendix
Table 12: Forecast Collapse for Width Schedule 1 at Depth 7
Comparison
Depth
Res. SR
FF SR
Δ SR
t
ΔROOS2
Δ FF5 α
Δ Turn.
ResNet vs NN
Shallow
2.40
2.31
0.09∗∗∗
6.46
0.01
0.13
0.55
Medium
2.46
2.29
0.17∗∗∗
4.22
0.01
0.08
0.63
Deep
2.42
2.13
0.29∗∗∗
3.74
0.02
0.04
1.17
ResNet+ vs NN+
Shallow
3.53
3.28
0.25∗∗
2.49
0.06
0.50
0.09
Medium
3.52
2.60
0.92∗∗∗
3.94
0.13
1.48
4.35
Deep
3.83
1.60
2.22∗∗∗
6.40
0.24
3.04
-12.15
Appendix
Table 13: Equal-Weighted Portfolio Performance and Sharpe Ratio Differences by Depth
Width
Depth
Model
Spec
ROOS2
Scale
VW SR
FF5 α
t(α)
Turnover
D1
D2
D3
D4
D5
D6
D7
D8
D9
D10
H − L
S1=[32]
1
ResNet+
resnetps1d1
0.33
0.75
2.11
3.79
7.77
268.70
-1.32
-0.22
0.21
0.50
0.56
0.77
0.90
1.15
1.51
2.45
3.77
S1=[32]
1
NN+
nnps1d1
0.33
0.75
2.11
3.79
7.77
268.70
-1.32
-0.22
0.21
0.50
0.56
0.77
0.90
1.15
1.51
2.45
3.77
S1=[32]
2
ResNet+
resnetps1d2
0.33
0.79
1.95
3.26
9.11
260.67
-1.37
-0.26
0.14
0.45
0.54
0.74
0.96
1.13
1.36
1.97
3.34
S1=[32]
2
NN+
nnps1d2
0.32
0.83
2.09
3.49
7.62
261.60
-1.38
-0.38
0.13
0.40
0.53
0.77
1.03
1.15
1.38
2.30
3.68
S1=[32]
3
ResNet+
resnetps1d3
0.36
0.75
1.85
3.33
7.81
246.43
-1.32
-0.27
0.31
0.53
0.61
0.80
0.98
1.03
1.29
1.94
3.26
S1=[32]
3
NN+
nnps1d3
0.32
0.85
1.98
3.37
7.44
265.04
-1.42
-0.33
0.11
0.41
0.53
0.78
1.01
1.14
1.38
2.16
3.58
Appendix
Table 14: Full ResNet+ and NN+ Value-Weighted Portfolio Sorts by Width and Depth
Width
Depth
Model
Spec
ROOS2
Scale
VW SR
FF5 α
t(α)
Turnover
D1
D2
D3
D4
D5
D6
D7
D8
D9
D10
H − L
S1=[32]
1
ResNet
resnets1d1
0.21
0.47
0.88
1.96
6.94
211.01
-0.49
0.10
0.49
0.47
0.68
0.69
0.79
0.87
0.94
1.26
1.75
S1=[32]
1
NN
nns1d1
0.21
0.47
0.88
1.96
6.94
211.01
-0.49
0.10
0.49
0.47
0.68
0.69
0.79
0.87
0.94
1.26
1.75
S1=[32]
2
ResNet
resnets1d2
0.21
0.47
0.87
1.93
6.83
210.68
-0.45
0.10
0.51
0.51
0.69
0.70
0.77
0.89
0.93
1.27
1.71
S1=[32]
2
NN
nns1d2
0.21
0.49
0.85
1.88
6.64
211.00
-0.46
0.11
0.45
0.55
0.66
0.70
0.80
0.87
0.93
1.23
1.69
S1=[32]
3
ResNet
resnets1d3
0.21
0.48
0.88
1.97
6.55
211.61
-0.47
0.11
0.50
0.45
0.72
0.67
0.78
0.88
0.95
1.28
1.75
S1=[32]
3
NN
nns1d3
0.20
0.49
0.87
1.86
6.58
215.51
-0.45
0.11
0.52
0.51
0.71
0.72
0.79
0.90
0.94
1.23
1.68
Appendix
Table 15: Full ResNet and NN Value-Weighted Portfolio Sorts by Width and Depth
Figure 8: Model Scale Reference for Asset Pricing Specifications. The horizontal axis uses a logarithmic scale. CAPM, the Fama–French model, and the Carhart model count regression parameters; Barra USE4 counts factors, and the factor zoo counts predictors. The remaining rows count neural-network parameters. The models are from Sharpe (1964) ; Fama and French (1993) ; Carhart (1997) ; MSCI (2013) ; Harvey et al. (2016) ; Devlin et al. (2019) ; Radford et al. (2023) ; Touvron et al. (2023) ; Hoffmann et al. (2022) ; Brown et al. (2020) ; Chowdhery et al. (2022)
We show that a deep neural network (DNN) trained to construct a stochastic discount factor (SDF) admits an additive decomposition separating nonlinear characteristic discovery from the pricing rule that aggregates them. This decomposition yields a linear factor representation governed by the Portfolio Tangent Kernel (PTK), which summarizes the network's learned features. In population, the implied SDF converges to a ridge-regularized version of the true SDF, with the degree of regularization determined by spectral complexity. Empirically, using U.S. equity data, the PTK representation delivers economically and statistically significant performance gains, while rising spectral complexity imposes tighter limits on finite-sample pricing.
Bryan Kelly, Boris Kuznetsov, Semyon Malamud +1
Yale School of Management, AQR Capital Management, and NBER · AQR Capital Management · NBER +5
Scaling network depth has been a central driver behind the success of modern foundation models, yet recent investigations suggest that deep layers are often underutilized. This paper revisits the default mechanism for deepening neural networks, namely residual connections, from an optimization perspective. Rigorous analysis proves that the layout of residual connections can fundamentally shape convergence behavior, and even induces an exponential gap in convergence rates. Prompted by this insight, we introduce adaptive neural connection reassignment (ANCRe), a principled and lightweight framework that parameterizes and learns residual connectivities from the data. ANCRe adaptively reassigns residual connections with negligible computational and memory overhead (<1%), while enabling more effective utilization of network depth. Extensive numerical tests across pre-training of large language models, diffusion models, and deep ResNets demonstrate consistently accelerated convergence, boosted performance, and enhanced depth efficiency over conventional residual connections.
Sparsity or complexity? In modern high-dimensional asset pricing, these are often viewed as competing principles: recent empirical evidence favors richer models, while economic intuition has long favored parsimony. We reconcile this tension by distinguishing capacity sparsity-restrictions on effective model capacity-from factor sparsity-the parsimonious structure of priced risks. Revisiting the benchmark empirical design of Didisheim et al. (2025), we combine nonlinear feature expansions with basis pursuit, using column generation and GPU acceleration to scale estimation to 432 million candidate factors. Reaching this scale reveals a reversal in out-of-sample performance: sparse portfolios trail dense ridgeless benchmarks at lower complexity but achieve a higher Sharpe ratio and lower pricing error at the largest candidate set. Capacity expansion and factor sparsity are therefore complements: enlarging the candidate space allows a parsimonious pricing kernel to outperform its dense counterpart.