Organizations: Research Institute for Interdisciplinary Sciences, School of Information Management and Engineering, Shanghai University of Finance and Economics, Shanghai 200433, China · Department of Decision Analytics and Operations, City University of Hong Kong, Hong Kong, China · Department of Industrial and Systems Engineering, University of Minnesota, Minneapolis, Minnesota 55455
Learn-then-differentiate (LTD) estimates gradients by fitting a model to simulation outputs and differentiating it. We develop a unified framework explaining what LTD differentiates and how accurately it estimates gradients. For models with a weighted representation, LTD differentiates a learned representation of the underlying probability measure. We then show how accuracy guarantees for fitted models translate into guarantees for gradients and higher-order derivatives, with rates approaching the standard Monte Carlo rate under suitable smoothness conditions. The framework recovers established results for kernel regression, local polynomial regression, and kernel ridge regression, and yields further guarantees for multiple kernel learning and smooth neural networks. These results provide a common foundation for understanding and analyzing LTD across learning methods.
Figures & tables
d
m
B
PW
LR
KR
LPR
KRR
MKL
NN
1
50
50k
0.366
2.565
3.764
3.818
3.461
3.623
3.164
1
50
500k
1.660
1.232
1.434
1.225
3.937
1
50
5m
0.974
0.614
0.412
0.342
3.698
1
200
50k
0.350
6.458
3.614
3.919
3.042
2.909
3.291
1
200
500k
1.582
1.372
1.335
1.195
6.142
1
200
5m
0.943
0.620
0.427
0.355
0.584
Table 1: Delta rRMSE (%) for the geometric Asian option example.
d
m
B
kernel PW
LR
KR
LPR
KRR
MKL
NN
1
50
50k
1.332
29.139
25.943
15.859
19.217
22.248
17.654
1
50
500k
15.499
8.882
7.347
5.365
11.477
1
50
5m
14.173
6.751
2.126
1.721
9.854
1
200
50k
1.224
100.641
22.541
14.718
12.493
11.806
16.239
1
200
500k
14.312
7.331
6.186
5.174
18.777
1
200
5m
12.719
5.685
2.229
1.752
3.529
Table 2: Diagonal Gamma rRMSE (%) for the geometric Asian option example.
d
B
FD
KR
LPR
KRR
MKL
NN
2
100k
11.162
24.791
8.203
8.605
9.416
13.775
2
1m
18.682
2.973
3.138
3.118
4.737
2
10m
17.827
1.309
1.262
1.369
1.938
4
100k
10.747
30.005
16.605
13.069
13.031
15.419
4
1m
26.928
10.703
7.540
7.696
9.055
4
10m
25.574
7.692
3.946
4.412
5.444
Table 3: Overall nRMSE (%) for the wireless communication network example.
d
Parameter
B
FD
KR
LPR
KRR
MKL
NN
2
p1
100k
11.956
24.330
9.003
9.419
10.673
14.875
2
p1
1m
19.361
3.425
3.681
3.585
5.072
2
p1
10m
17.153
1.381
1.437
1.565
2.141
2
p2
100k
10.508
25.142
7.528
7.920
8.317
12.861
2
p2
1m
18.140
2.569
2.644
2.703
4.462
2
p2
10m
18.331
1.250
1.107
1.195
1.767
Table 4: Componentwise rRMSE (%) for the wireless communication network example.
d
B
Global
KR
LPR
1
50,000
50
834
834
1
500,000
500
1077
1159
1
5,000,000
1000
1392
1610
2
50,000
50
324
289
2
500,000
500
484
529
2
5,000,000
1000
784
900
Table K.1: Training-site allocations for the option example. Global counts apply to KRR, MKL, and NN.
Quantity
Symbol
Value
Region half-width
L
500m
Antenna height
h1=h2
30m
Peak antenna gain
Gmax
17dBi
Horizontal 3-dB beamwidth
ϕ3dB
65∘
Vertical 3-dB beamwidth
ψ3dB
10∘
Maximum pattern attenuation
Amax
30dB
Table L.2: Fixed parameters in the wireless simulator.
d
B
Global
KR
LPR
2
100,000
200
121
121
2
1,000,000
200
192
216
2
10,000,000
200
304
383
4
100,000
1000
625
625
4
1,000,000
1000
1347
1570
4
10,000,000
1000
2901
3944
Table L.3: Training-site allocations for the wireless example. Global counts apply to KRR, MKL, and NN.
Carnegie Mellon University, Pittsburgh, USA · Massachusetts Institute of Technology, Cambridge, USA · Chalmers University of Technology & University of Gothenburg, Gothenburg, Sweden +1