Deep hedging learns trading policies from historical or simulated market trajectories, yet under nonstationarity these training paths may not represent future market conditions. We propose WRAP (Wasserstein-Reweighting Adversarial Perturbation), a drift-aware adversarial training framework derived from a two-budget distributionally robust optimization (DRO) formulation. The formulation is anchored to a weighted empirical reference distribution whose fixed baseline weights are chosen to balance sampling uncertainty against temporal drift. Around this reference distribution, the ambiguity set addresses two complementary forms of distributional misspecification by allowing an adversary to reweight the observed trajectories subject to a φ-divergence constraint and perturb their paths subject to an optimal-transport (OT) constraint. We derive a joint first-order expansion in which the leading-order increase over the nominal expected loss decomposes into a reweighting contribution determined by the dispersion of hedging losses across trajectories and a transport contribution determined by the sensitivity of the loss to path perturbations. This expansion yields an explicit finite-dimensional adversarial attack that replaces the distributional inner supremum with a tractable first-order approximation. Across stationary and nonstationary Heston dynamics and a generalized affine diffusion (GAD), the experiments show complementary benefits from reweighting and transport, with joint adversarial training providing the largest gains under nonstationarity.
Figures & tables
Algorithm 1 W asserstein R eweighting A dversarial P erturbation ( WRAP )
Method
ε
τ
Val. CVaR (↓)
Near CVaR (↓)
Far CVaR (↓)
Far Mean PnL (↑)
Far Turnover (↓)
ERM-U
–
–
6.868±0.028
7.952±0.137
8.992±0.236
−4.528±0.055
10.450±0.955
ϕ -U
–
1.00
6.324±0.034
7.239±0.118
8.128±0.248
−4.557±0.038
10.717±1.774
OT-U
0.20
–
6.245±0.093
7.050±0.180
7.811±0.397
−4.594±0.044
10.056±1.572
WRAP-U
0.30
0.60
6.005±0.058
6.641±0.136
7.268±0.193
−4.571±0.029
10.437±2.500
ERM-W
–
–
6.653±0.044
7.735±0.126
8.707±0.249
−4.532±0.042
7.995±0.650
ϕ -W
–
1.20
6.259±0.036
7.097±0.169
7.919±0.305
−4.553±0.042
8.033±0.455
Table 1 : Results for the nonstationary Heston model.
Figure 1 : Loss landscapes across deep-hedging training schemes from Table 1 : (A) ERM-W , (B) ϕ -W , (C) OT-W , and (D) WRAP-W . Each surface depicts the mean hedging loss of a separately trained network on the same reference trajectories Xn . The axes show terminal moneyness and the root mean square (RMS) of daily log returns in percent. Markers represent reference paths and adversarial trajectories Xn , with sizes proportional to their respective probabilities wn and wn . Arrows indicate selected transport displacements. Within each main panel, subpanels 1–4 report the full feature range, the moneyness marginal, the log-return marginal, and the loss landscape.
Figure 2 : Pathwise effects in nonstationary Heston: (A) ERM-U , (B) ϕ -U , (C) WRAP-U , and (D) magnified view of the original and transported paths highlighted in (C). Panels (A)–(C) show sampled trajectories and highlight the highest-weight and most-displaced paths.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Algorithm 2 Multistep W asserstein R eweighting A dversarial P erturbation ( WRAP )
Parameter
Heston
GAD
Architecture and Optimization
Base architecture
Per-date MLP hedger
Hidden units
20
Hidden layers per date
2
Activation
ReLU
Normalization
Batch normalization
Appendix
Table 2: Hyperparameter settings for the Heston and GAD experiments.
Method
ε
τ
Val. CVaR (↓)
Test CVaR (↓)
Mean PnL (↑)
Turnover (↓)
ERM-U
–
–
3.634±0.013
3.629±0.018
−2.199±0.005
3.084±0.124
ϕ -U
–
0.25
3.627±0.014
3.629±0.025
−2.198±0.004
3.283±0.266
OT-U
0.10
–
3.550±0.011
3.543±0.016
−2.199±0.004
3.156±0.064
WRAP-U
0.10
–
3.550±0.011
3.543±0.016
−2.199±0.004
3.156±0.064
Appendix
Table 3 : Results for the stationary Heston model with the CVaR objective.
Method
ε
τ
Val. CVaR (↓)
Test CVaR (↓)
Mean PnL (↑)
Total TC (↓)
ERM-U
–
–
4.500±0.011
4.494±0.016
−3.003±0.006
0.803±0.004
ϕ -U
–
0.05
4.491±0.009
4.485±0.015
−3.019±0.004
0.819±0.004
OT-U
0.05
–
4.488±0.010
4.484±0.013
−3.002±0.010
0.802±0.008
WRAP-U
0.05
0.05
4.487±0.009
4.481±0.012
−3.014±0.006
0.815±0.007
Appendix
Table 4 : Results for the stationary Heston model with transaction costs and the CVaR objective.
Figure 3 : Stationary Heston validation grids with uniform reference weights and the CVaR objective, corresponding to Tables 3 and 4 . The surface reports the percentage reduction in validation CVaR relative to the corresponding ERM-U baseline. Crosses indicate the evaluated budget pairs, with contours interpolating between them; the selected WRAP budgets are (ε,τ)=(0.10,0) in (a) and (0.05,0.05) in (b).
Method
ε
τ
Val. M–CVaR (↓)
Test M–CVaR (↓)
Test CVaR (↓)
Mean PnL (↑)
Total TC (↓)
ERM-U
–
–
3.717±0.010
3.714±0.008
4.500±0.014
−2.927±0.004
0.727±0.003
ϕ -U
–
0.05
3.716±0.007
3.713±0.009
4.486±0.014
−2.940±0.004
0.740±0.002
OT-U
0.05
–
3.708±0.009
3.706±0.007
4.502±0.013
−2.909±0.004
0.709±0.003
WRAP-U
0.05
–
3.708±0.009
3.706±0.007
4.502±0.013
−2.909±0.004
0.709±0.003
Appendix
Table 5 : Results for the stationary Heston model with transaction costs and the M–CVaR objective.
Figure 4 : Nonstationary Heston validation grids for Table 1 . Colours show percentage reductions Val. CVaR relative to ERM-U in (a) and ERM-W in (b); Black crosses mark the budget pairs in ( 47 ) ; contours interpolate between them. Red crosses mark the selected WRAP budgets: (ε,τ)=(0.30,0.60) in (a) and (0.40,1.00) in (b)
Method
ε
τ
Val. CVaR (↓)
Near CVaR (↓)
Far CVaR (↓)
Far Mean PnL (↑)
Far Total TC (↓)
ERM-U
–
–
7.592±0.077
8.250±0.155
8.935±0.260
−5.562±0.078
1.002±0.073
ϕ -U
–
1.20
7.122±0.042
7.778±0.131
8.434±0.225
−5.714±0.082
1.163±0.076
OT-U
0.10
–
7.752±0.073
8.456±0.134
9.170±0.195
−5.565±0.101
0.976±0.095
WRAP-U
0.10
0.80
7.082±0.041
7.691±0.111
8.349±0.189
−5.695±0.071
1.130±0.065
Appendix
Table 6 : Results for the nonstationary Heston model with transaction costs and the CVaR objective.
Method
ε
τ
Val. M–CVaR (↓)
Near M–CVaR (↓)
Far M–CVaR (↓)
Far Mean PnL (↑)
Far Total TC (↓)
ERM-U
–
–
6.281±0.048
6.766±0.077
7.251±0.149
−5.552±0.084
0.994±0.077
ϕ -U
–
1.00
6.078±0.011
6.581±0.068
7.090±0.149
−5.721±0.089
1.170±0.079
OT-U
0.10
–
6.388±0.074
6.898±0.078
7.401±0.151
−5.562±0.122
0.974±0.112
WRAP-U
0.20
0.80
6.045±0.013
6.523±0.062
7.019±0.148
−5.681±0.098
1.112±0.092
Appendix
Table 7 : Results for the nonstationary Heston model with transaction costs and the M–CVaR objective.
Figure 5 : Stock-hedging policies for stationary Heston with uniform reference weights and the CVaR objective, using the selections in Table 3 : (A) ERM-U , (B) ϕ -U , (C) OT-U , and (D) WRAP-U . Colours show stock holdings averaged over validation paths and five seeds within each date–moneyness bin. Dashed curves mark the 5th and 95th moneyness percentiles; the solid line marks the strike.
Figure 6 : Stock-hedging policies for stationary Heston with uniform reference weights, the CVaR objective, and proportional transaction costs c=0.005 , using the selections in Table 4 . Panels, averaging, and plotting conventions follow Figure 5 , with the same colour scale.
Figure 7 : Stock-hedging policies for nonstationary Heston with uniform reference weights and the CVaR objective, using the selections in Table 1 : (A) ERM-U , (B) ϕ -U , (C) OT-U , and (D) WRAP-U . Averaging and plotting conventions follow Figure 5 , with the same colour scale.
Figure 8 : Stock-hedging policies for nonstationary Heston with drift-aware reference weights and the CVaR objective, using the selections in Table 1 : (A) ERM-W , (B) ϕ -W , (C) OT-W , and (D) WRAP-W . The displayed ratio r=ϱ/εϱ≈0.00853 controls reference weighting. Averaging and plotting conventions follow Figure 5 , with the same colour scale.
Stock
Split
μ (%)
Median (%)
σ (%)
MaxDD (%)
AC(1)
pKS
AAPL
Train
30.24
25.39
29.33
−43.80
−0.000
Val
29.90
24.90
33.55
−31.43
−0.133
0.016
Test
23.51
29.92
27.87
−33.36
0.045
0.046
MSFT
Train
21.56
16.11
26.86
−43.95
−0.084
Val
27.58
24.65
32.60
−37.15
−0.197
<0.001
Test
16.54
28.00
22.16
−23.73
−0.018
0.682
Appendix
Table 8: Descriptive statistics of daily simple adjusted-close returns across the training, validation, and test periods. Mean, median, volatility, and maximum drawdown are reported in percentages. The pKS column reports two-sample Kolmogorov–Smirnov p -values relative to the corresponding training-period returns.
Training
Stock
Method
ε
τ
Val. OCE (↓)
Test OCE (↓)
Mean PnL (↑)
Turnover (↓)
Fix
AAPL
ERM-U
–
–
0.287±0.002
0.258±0.001
−0.227±0.000
0.072±0.002
ϕ -U
–
0.05
0.308±0.012
0.272±0.008
−0.227±0.002
0.089±0.005
OT-U
0.025
–
0.299±0.003
0.267±0.002
−0.226±0.001
0.072±0.005
WRAP-U
0.05
0.15
0.289±0.002
0.261±0.001
−0.226±0.001
0.066±0.001
Fix
AMZN
ERM-U
–
–
0.310±0.001
0.289±0.001
−0.255±0.000
0.069±0.001
ϕ -U
–
0.05
0.328±0.004
0.301±0.003
−0.252±0.002
0.083±0.003
Appendix
Table 9 : Hedging under GAD dynamics calibrated to successive periods of real equity prices.
Attack
Method
ε
τ
Val. CVaR (↓)
Test CVaR (↓)
Mean PnL (↑)
Turnover (↓)
Baseline
ERM-U
–
–
3.634±0.013
3.629±0.018
−2.199±0.005
3.084±0.124
One-step WRAP
ϕ -U
–
0.25
3.627±0.014
3.629±0.025
−2.198±0.004
3.283±0.266
OT-U
0.10
–
3.550±0.011
3.543±0.016
−2.199±0.004
3.156±0.064
WRAP-U
0.10
–
3.550±0.011
3.543±0.016
−2.199±0.004
3.156±0.064
Multistep-WRAP
ϕ -U
–
0.25
3.627±0.014
3.629±0.025
−2.198±0.004
3.283±0.266
OT-U
0.10
–
3.552±0.011
3.545±0.015
−2.200±0.004
3.167±0.060
Appendix
Table 10 : Results for one-step and multistep WRAP in the stationary Heston setting.
Figure 9 : Pathwise effects in nonstationary Heston: (A) ERM-U , (B) ϕ -U , (C) WRAP-U , and (D) magnified view of the original and transported paths highlighted in (C). Panels (A)–(C) show sampled trajectories and highlight the highest-weight and most-displaced paths.