Targeted search shows that random-device testing underestimates worst-case error in a simulated wave-based neural operator
Authors: Samrendra Roy, Jason Yoo, Souvik Chakraborty, Syed Bahauddin Alam
Organizations: Department of Nuclear, Plasma, and Radiological Engineering, University of Illinois Urbana-Champaign, Urbana, IL 61801, USA · Department of Applied Mechanics, Indian Institute of Technology Delhi, Hauz Khas, New Delhi 110016, India · Yardi School of Artificial Intelligence (ScAI), Indian Institute of Technology Delhi, Hauz Khas, New Delhi 110016, India · National Center for Supercomputing Applications (NCSA), University of Illinois Urbana-Champaign, Urbana, IL 61801, USA
Wave-based processors promise fast, energy-efficient Fourier layers for neural operators. They are usually validated on randomly sampled devices, but using them requires knowing how large their error can become under fabrication and alignment variation. In a stylised numerical case study, a hybrid Fourier neural operator runs its four spectral layers on simulated coherent 4f processors with 32 toleranced knobs, whose half-widths are representative rather than calibrated. For 120 models (four tasks, six training methods, five seeds), we compared the worst of N random in-spec devices with a searched one. On a deterministic simulator with one frozen draw of the random static errors, the searched device's held-out error was 1.08-3.10 times the maximum over 200 Monte Carlo devices and 1.06-2.71 times that over 1000. With 20 fresh static draws, it still exceeded the maximum over 200 random devices in 116 of 120 models. Under uniform sampling, the probability of drawing such a device is at most 0.37% per model (two-sided 95% Clopper-Pearson), which says nothing about how large its error is. The gap persisted with uniform or Sobol' sampling at the search's budget, shared knobs, a second crosstalk model, box scales of 0.25-2 and a pixel-level device model. Models trained only with random static errors reached 3.7-39.9 times their nominal error on searched devices, and fine-tuning on random and gradient-searched devices gave the lowest searched error of the six in all 20 task-seed pairs. For two heat-exchanger quantities, a search targeted at each exceeded the worst of 1000 random devices in all 39 models, and hence the Wilks 95/95 limit (worst of 59). For the mean pressure of 11 models, no random device exceeded a 1% error threshold, but the searched device did. Random testing estimates how often errors exceed a threshold; worst-device search gives a lower bound on how large they can be.
Figures & tables
Figure 1: Simulated hybrid wave-based neural operator, tolerance box and testing approaches. a , Model pipeline with tensor shapes (channels × height × width). Blue marks the wave-based (4f optical) spectral paths Kl and amber the digital parts. Darcy and Navier–Stokes (NS) take an input field. The inputs of the lid-driven cavity (LDC; 90-dimensional lid-history encoding) and the heat exchanger (HX; 102-dimensional condition vector) pass through a digital multilayer perceptron (MLP) to a coarse latent grid, which is upsampled bilinearly. Two grid coordinates are appended before the lift, and GELU follows the first three hybrid layers. HX reads its output at the 3977 mesh nodes with a digital readout. b , One hybrid layer: a simulated coherent 4f processor in parallel with a digital 1×1 convolution, with an idealised phase-sensitive camera readout against a holographic reference beam. Red circles mark the eight toleranced knobs per layer where they act; they are independent across layers in the primary 32-D box. The table gives the representative half-widths. The two lines at the bottom of the panel list the random static errors, present in both configurations, and the converters and detection noise, which the deterministic configuration ( det ) omits. c , Random testing samples Monte Carlo (MC) devices uniformly from the box, whereas targeted search seeks a high-error device, whose error is a lower bound on the worst case. Random corners belong to the unseen attacks of the search, not to random testing. The error surface is illustrative.
Figure 2: Search finds in-spec devices that random testing misses. a , Primary endpoint RN (equation ( 1 )) for all 120 models on det and held-out samples, for N=200 (filled) and N=1000 (open), five seeds per arm. All values exceed 1. LDC errors are effectively pressure errors (Methods). b , Development errors of 100,000 uniformly drawn devices for the Box seed-0 model of each task (histogram, logarithmic counts). The dotted and dashed lines mark the highest error among the first 1000 devices and among the first 16,256, the budget of the search, and the red line marks the searched device. No device reached it (Section 2.3 ). c , Highest development error reached so far, divided by the highest found over all methods, after 64, 512 and all evaluations, mean over the five box-trained arms. The curves show budget progress, not convergence. DE and CMA-ES (seed 0) are counted in forward evaluations, and the PGD curve covers only its four independent starts (124 forward passes, most with a backward pass; Methods and Supplementary Note 2).
Figure 3: The gap under other tolerance sets and crosstalk models, and the searched devices. a , R200 against the box scale κ (all half-widths scaled) for the seed-0 Box, Smooth-mask box and Mixed-PGD FT models on Darcy, NS and LDC (Supplementary Table S9). b , Box arm only, same sweep: searched error, MC maximum of 200 and MC mean, each divided by the nominal error. c , R200 in the shared-knob 8-D box (the same eight knobs in all layers) against the 32-D box for the same seed-0 models (Supplementary Table S8). Colour gives the arm and marker shape the task (legend at the top). The dashed line marks equality, and grey rings mark models whose highest-error device was a box corner. d , R200 with circular-phase crosstalk against field crosstalk (seeds 0–2; Supplementary Table S10). The circular-phase models were retrained, so each point compares complete training-and-simulator configurations rather than isolating the crosstalk law. e , Searched devices zsearch (Box and Mixed-PGD FT, seed 0) in normalised coordinates, where, for knobs 6 and 7, z=−1 is the sharpest crosstalk and the lowest extinction ratio. Each row is one model.
Figure 4: Training arms on four tasks. a–d , Nominal device (open), maximum over 200 MC devices (light) and searched device (filled) for each arm, all on det and held-out samples. Each symbol is the mean over five seeds, with a vertical line from the seed minimum to the seed maximum. The dashed line is the digital FNO on ideal hardware (three seeds), and the y -axes are logarithmic. Supplementary Tables S1 and S2 list all models. e , Change of the searched (filled) and unseen-attack (open) error of Mixed-PGD FT relative to Box FT with the same seed. Supplementary Table S11 lists the paired values.
Figure 5: Heat exchanger (3977-node mesh). a , Per-channel relative L2 error (pressure p , in-plane z - and y -velocities u1 and u2 ) on det for three arms (Random-error-aware, Box and Mixed-PGD FT). Devices: nominal, the MC device selected by development error from all 1000 (not the held-out MC maximum used in RN ) and the searched device. Values are means over five seeds, with vertical lines from the minimum to the maximum over seeds. Black marks the digital FNO (three seeds), and the y -axis is logarithmic. Supplementary Table S12 lists the per-channel and joint errors. b , Box seed 0, held-out sample with the median nominal error: true fields and the errors (prediction minus true field) on the nominal device, on the worst of the first 200 MC devices (selected by development error) and on the searched device. The error maps of a channel share a symmetric colour scale, clipped at the 99.5th percentile of the absolute error, and true fields span their 0.5th to 99.5th percentiles. Error colour bars use the row label’s units. Each map gives that channel’s relative L2 error (primary metric 0.31%, 0.40% and 1.12%).
Task
Arm
Nominal
MC max 200
Unseen-attack best
Searched (mean ± s.d.)
R200
Δ vs Box FT (%)
Darcy
Random-error-aware
1.58
7.58
10.6
10.7 ± 2.0
1.35–1.49
+92 [+51, +142]
Box
2.46
2.95
5.27
5.96 ± 0.18
1.88–2.14
+7 [ − 1, +12]
Smooth-mask box
3.02
3.62
6.09
6.99 ± 0.27
1.86–2.03
+26 [+17, +34]
Box FT
2.31
2.76
4.84
5.57 ± 0.32
1.85–2.12
–
Mixed-PGD FT
2.03
2.45
4.02
4.35 ± 0.19
1.71–1.88
− 22 [ − 26, − 18]
SAM FT
2.20
2.96
4.81
5.09 ± 0.32
1.62–1.93
− 9 [ − 12, − 4]
Table 1: Training arms ( det , held-out samples, relative L2 error in %; HX: mean of the per-channel errors; LDC: joint error over p , k and ∣u∣ , dominated by p ; see Methods). Values are means over five seeds (digital FNO: three seeds). Unseen-attack best: error of the highest-error device found by corners, CA or CMA-ES. Searched: mean ± s.d. over seeds. R200 : range over seeds. Δ : change of the searched error relative to Box FT with the same seed, mean [min, max]. Digital FNO: ideal hardware.
Task
Arm, seed
nominal
MC mean
max 200
max 1000
unseen
DE/PGD
searched
R200
R1000
method
Darcy
Random-error-aware s0
1.58
5.07
7.28
7.72
10.0
10.0
10.0
1.38
1.30
Pd
Darcy
Random-error-aware s1
1.63
4.25
6.23
6.38
9.00
9.00
9.00
1.44
1.41
A
Darcy
Random-error-aware s2
1.58
4.55
6.69
7.19
9.65
9.65
9.65
1.44
1.34
A
Darcy
Random-error-aware s3
1.60
5.02
7.25
8.25
10.5
10.8
10.8
1.49
1.31
Pd
Darcy
Random-error-aware s4
1.53
6.90
10.5
11.4
13.7
14.1
14.1
1.35
1.23
Pd
Darcy
Box s0
2.46
2.03
2.90
3.24
5.09
5.77
5.77
1.99
1.78
Pc
Table S1: Primary endpoint, Darcy and NS ( det , static draw 0, held-out samples, error in %). MC: 1000 devices, maximum over the first 200 and over all 1000. Unseen: best of corners, CA and CMA-ES. DE/PGD: best of DE and PGD. Searched: the searched device zsearch , the device with the highest development error over all searches. Method identifies the search that found it: A = CA, K = corners, Pd/Pc/Pa/Pk/Pm/Pr = PGD from DE, jittered centre, CA, best corner, best MC device or random start, C0, C1, … = CMA-ES runs. No searched device in these tables came from DE alone.
Task
Arm, seed
nominal
MC mean
max 200
max 1000
unseen
DE/PGD
searched
R200
R1000
method
LDC
Random-error-aware s0
1.08
13.6
26.3
26.3
35.8
35.8
35.8
1.36
1.36
A
LDC
Random-error-aware s1
0.702
10.4
20.1
20.7
27.6
28.1
28.1
1.40
1.35
Pa
LDC
Random-error-aware s2
1.02
9.07
17.8
18.9
24.2
24.6
24.6
1.38
1.30
Pa
LDC
Random-error-aware s3
1.64
9.72
18.8
21.1
28.7
28.7
28.7
1.52
1.36
A
LDC
Random-error-aware s4
0.948
11.9
21.8
23.5
33.5
33.5
33.5
1.53
1.43
A
LDC
Box s0
0.820
0.860
1.17
1.36
2.55
2.60
2.60
2.22
1.91
Pd
Table S2: Primary endpoint, LDC and HX , defined as in Supplementary Table S1 . HX error is the mean of the per-channel errors.
real device (%)
det : devices of 1000 exceeding c× nominal
Task
Arm, seed
nominal
max 200
zsearch mean
zsearch max
nested
c=1.1
1.25
1.5
2
UCB q0.99
Darcy
Random-error-aware s0
1.62
7.12
10.2
10.3
10.2
1000
1000
998
964
7.48
Darcy
Random-error-aware s1
1.66
6.36
8.97
9.06
8.98
1000
1000
997
899
6.26
Darcy
Random-error-aware s2
1.65
6.91
9.69
9.76
9.68
1000
1000
997
916
6.79
Darcy
Random-error-aware s3
1.68
7.77
10.9
11.2
10.9
1000
1000
999
948
7.95
Darcy
Random-error-aware s4
1.58
11.3
14.4
15.1
14.7
1000
1000
1000
974
10.7
Table S3: Secondary endpoint and reliability, Darcy and NS. real configuration (%): nominal error (20 static draws), maximum of 200 MC devices with independent static draws, and zsearch re-scored on 20 static draws (mean and maximum). Nested: PGD over eight common static draws, re-scored on 20 fresh draws. Reliability on det : number of the 1000 MC devices with error above c times nominal, and the 95% upper confidence bound on the 0.99 quantile (%).
real device (%)
det : devices of 1000 exceeding c× nominal
Task
Arm, seed
nominal
max 200
zsearch mean
zsearch max
nested
c=1.1
1.25
1.5
2
UCB q0.99
LDC
Random-error-aware s0
1.20
25.4
35.3
35.9
35.4
995
989
987
975
25.6
LDC
Random-error-aware s1
1.26
19.8
28.2
29.3
28.0
1000
1000
991
981
19.9
LDC
Random-error-aware s2
1.29
17.2
25.2
26.0
25.2
991
987
971
953
17.8
LDC
Random-error-aware s3
1.25
20.0
28.5
29.2
28.5
972
966
958
925
19.4
LDC
Random-error-aware s4
1.25
21.9
33.7
34.3
33.8
994
989
983
969
22.2
Table S4: Secondary endpoint and reliability, LDC and HX , as in Supplementary Table S3 .
uniform
Sobol’
corners
corners + uniform
Task
Arm (s0)
searched
error
Rmatch
error
Rmatch
error
Rmatch
error
Rmatch
Darcy
Random-error-aware
10.0
8.30
1.21
8.24
1.22
10.0
1.00
9.83
1.02
Box
5.77
3.52
1.64
3.73
1.55
4.75
1.21
4.75
1.21
Smooth-mask box
6.64
4.11
1.62
3.96
1.68
5.13
1.30
5.11
1.30
Box FT
5.17
2.96
1.74
3.35
1.54
4.62
1.12
4.62
1.12
Mixed-PGD FT
4.05
2.63
1.54
2.69
1.50
3.90
1.04
3.90
1.04
Table S5: Budget-matched samplers (seed-0 models, mode-level device, det , static draw 0, held-out samples, error in %). Each sampler draws Nmatch=16,256 devices and selects the one with the highest development error. Error: held-out error of that device. Rmatch : searched error divided by it. Corners + uniform: the mixed sampler, which alternates corner and uniform devices (Supplementary Note 3).
best development error (%)
generalised Pareto tail
Task (Box s0)
103
104
105
searched
≥ searched
R at 105
ξ
N∗ ( q=0.98 )
N∗ (0.99)
N∗ (0.995)
bootstrap 5%
Darcy
3.03
3.27
3.41
5.41
0
1.59
−0.030
3.8×1016
3.6×1017
5.4×1011
9.5×109
NS
16.7
17.4
17.9
26.2
0
1.45
−0.080
∞
∞
5.0×1017
8.2×1019
LDC
1.28
1.39
1.72
2.54
0
1.45
−0.104
∞
∞
∞
2.9×1014
HX
0.526
0.587
0.599
1.27
0
2.12
+0.010
∞
9.0×1015
2.0×1010
7.0×1010
Table S6: 100,000 uniform devices and generalised Pareto tail fits (Box seed 0, mode-level device, det , development errors in %). Best development error among the first 103 , 104 and 105 devices; searched: development error of zsearch ; ≥ searched: devices of 105 that reached it; R at 105 : held-out error of zsearch divided by that of the device selected from 105 . Tail: shape ξ (threshold at the 0.99 quantile), N∗ at three thresholds (infinite when the fitted tail ends below the searched error) and the 5th percentile of N∗ over 200 bootstrap resamples (0.99 quantile). The tail values are extrapolations of a fitted law (Supplementary Note 3).
Rmatch
Model
Arm, seed
nominal
MC mean
max 200
max 1000
searched
R200
uniform
corners
mixed
nom./MC mean
searched device
HX
Random-error-aware s0
0.262
1.01
1.47
1.57
2.58
1.75
1.48
1.06
1.09
0.26
0/0/0
Random-error-aware s1
0.263
1.04
1.49
1.64
2.68
1.80
1.56
1.09
1.09
0.25
0/0/0
Random-error-aware s2
0.245
1.15
1.59
1.77
2.74
1.73
1.42
1.05
1.05
0.21
0/0/0
Box s0
0.510
0.332
0.493
0.634
1.11
2.25
1.68
1.44
1.44
1.54
4/4/8
Box s1
0.558
0.352
0.515
0.671
1.15
2.23
1.60
1.33
1.34
1.58
4/4/6
Table S7: Pixel-level device ( det , static draw 0, held-out samples, error in %; HX: mean of the per-channel errors). HX and Darcy: retrained models with the cell-integral shift; Darcy, sinc: the same arms and seeds retrained with the band-limited shift; Transfer: seed-0 models trained on the mode-level device and evaluated on the pixel-level device (cell integral). Nominal: the nominal device (0.5-pixel fringing, 30 dB), not the box centre. Rmatch : as in Supplementary Table S5 for the uniform, corner and mixed samplers. Nom./MC mean: nominal error divided by the MC mean. Searched device: layers at the sharpest fringing / layers at the lowest extinction ratio / shift coordinates within 0.05 of zero (of 8).
Task
Arm (s0)
nominal
MC mean
max 200
max 1000
all 256 corners
searched
R200
R1000
method
Darcy
Random-error-aware
1.58
5.05
8.55
8.77
9.85
9.85
1.15
1.12
K
Darcy
Box
2.46
2.38
4.30
4.62
4.75
5.88
1.37
1.27
Pd
Darcy
Smooth-mask box
2.85
3.01
4.65
5.10
5.18
6.50
1.40
1.27
Pd
Darcy
Box FT
2.27
2.25
3.88
4.13
4.65
5.25
1.35
1.27
Pm
Darcy
Mixed-PGD FT
2.00
2.16
3.16
3.26
3.95
4.12
1.31
1.26
Pd
Darcy
SAM FT
2.15
2.43
3.92
4.07
4.83
4.94
1.26
1.22
Pd
Table S8: Shared-knob 8-D box (seed 0, the same eight knobs in all four layers, det , held-out samples, error in %). Method codes as in Supplementary Table S1 .
Task
Arm (s0)
κ
MC mean
max 200
corners
unseen
searched
R200
Darcy
Box
0.25
2.11
2.56
2.20
2.26
3.17
1.24
0.5
1.90
2.66
2.10
2.26
3.95
1.48
1
2.01
2.90
4.27
4.88
5.77
1.99
1.5
2.65
4.31
6.75
7.52
7.52
1.75
2
3.58
5.98
9.04
9.91
9.91
1.66
Darcy
Mixed-PGD FT
0.25
1.84
2.05
1.89
1.90
2.38
1.16
Table S9: Tolerance sweep (seed 0, half-widths scaled by κ , reduced protocol, det , held-out samples, error in %). Nominal errors are as in Supplementary Tables S1 and S2 .
Task
Arm, seed
nominal
MC mean
max 200
max 1000
unseen
searched
R200
method
Darcy
Random-error-aware s0
1.86
7.07
9.77
10.8
14.1
14.7
1.50
Pa
Darcy
Random-error-aware s1
1.91
5.51
7.27
7.92
10.6
10.6
1.45
Pa
Darcy
Random-error-aware s2
1.86
6.05
8.38
8.93
12.7
12.9
1.54
Pa
Darcy
Box s0
3.16
2.05
2.45
2.82
4.36
4.38
1.78
Pa
Darcy
Box s1
3.50
2.08
2.67
3.02
4.34
4.38
1.64
Pa
Darcy
Box s2
3.75
2.15
3.20
3.21
4.60
5.32
1.66
Pc
Table S10: Circular-phase crosstalk set. Retrained models, seeds 0–2, with a reduced search of one DE seed and one CMA-ES seed from the centre ( det , held-out samples, error in %). Unseen: best of the unseen attacks.
searched (%)
unseen attacks (%)
Task
seed
Box FT
Mixed-PGD FT
change
Box FT
Mixed-PGD FT
change
Darcy
s0
5.17
4.05
− 22%
4.76
3.93
− 17%
s1
5.95
4.43
− 26%
5.33
4.07
− 24%
s2
5.37
4.43
− 18%
4.75
4.05
− 15%
s3
5.53
4.55
− 18%
4.71
4.03
− 15%
s4
5.82
4.29
− 26%
4.67
4.03
− 14%
Table S11: Paired values behind main-text Fig. 4e. Searched error and best unseen-attack error of Box FT and Mixed-PGD FT for every task and seed, and the change of Mixed-PGD FT relative to Box FT. All values are on det and held-out samples (%), and the HX error is the mean of the per-channel errors.
Arm
Device
primary (mean of channels)
p
u1
u2
joint median
Random-error-aware
nominal
0.243
0.314
0.209
0.206
0.275
MC (dev-selected)
3.94
5.65
3.14
3.04
5.66
searched
4.82
6.87
3.87
3.72
6.88
Box
nominal
0.351
0.449
0.305
0.299
0.406
MC (dev-selected)
0.501
0.636
0.432
0.436
0.576
searched
1.18
1.46
1.06
1.01
1.31
Table S12: Heat exchanger: per-channel and joint errors (%, mean over seeds, held-out samples, det ). MC (dev-selected): the MC device selected by development error from all 1000 MC devices. Joint median: joint relative L2 error over all channels, median over samples.
Model
Knobs
κ
expansion-point error (%)
worst found (%)
bound (%)
bound / worst
float slack
Box
32
0.001
2.396
2.400
2.416
1.01
–
Box
32
0.003
2.393
2.405
2.448
1.02
–
Box
32
0.005
2.391
2.410
2.578
1.07
–
Box
32
0.007
2.388
2.416
2.978
1.23
–
Box
32
0.01
2.385
2.424
4.915
2.03
–
Box
32
0.03
2.360
2.478
315.7
127
–
Table S13: Certificate bound across box scales κ (v2 simulator, Darcy seed 0, det , static draw 0, mean over ten test inputs). Expansion-point error: error at the Taylor model’s expansion point, which is not the nominal device and moves with κ (Supplementary Note 6). Worst found: per-input search, floored by the sampled devices.
Effect alone
passive s0
passive s1
ideal (calibrated)
mean ratio
none (ideal passive)
1.64
1.73
2.04
1.00
input DAC, 8 bits
1.65
1.73
2.05
1.00
mask phase, 8 bits
1.65
1.73
2.10
1.01
mask amplitude, 8 bits
1.65
1.75
2.05
1.00
output ADC, 10 bits
1.64
1.73
2.04
1.00
extinction ratio, 30 dB
8.78
8.03
12.04
5.29
Table S14: v1 pilot effect decomposition (legacy simulator): each non-ideality alone at its magnitude in the v1 realistic configuration, the v1 counterpart of real (error in %, random effects averaged over 20 frozen devices).
Setting
Frozen learned library (%)
Frozen random optics (%)
calibration chip
other chips
penalty
calibration chip
other chips
penalty
main ( Ntrain=1000 )
20.35
21.50
1.16
24.12
25.88
1.76
pre-mix adapter
16.57
18.63
2.06
19.60
22.03
2.43
Ntrain=250
25.07
25.81
0.74
27.48
29.99
2.51
Ntrain=100
27.83
28.46
0.64
30.17
32.62
2.46
slim adapter
25.51
28.12
2.62
24.15
26.25
2.10
Table S15: Transfer of adapters calibrated on one simulated chip to other simulated chips (frozen-library pilot, v1 simulator, NS unless noted, relative L2 error in %, penalty in percentage points; Ntrain : number of training samples). Pre-mix adapter: the adapter adds a digital pre-mix before each optical pass; slim adapter: a reduced adapter with a diagonal skip and a 32-wide projection; 16 blocks: 16 optical blocks, four per layer, instead of one shared bank of four; phase-only: phase-only masks.
pressure p
TKE k
speed ∣u∣
Arm
nom.
MC
search
nom.
MC
search
nom.
MC
search
Random-error-aware
1.08
21.0
30.1
10.7
31.4
43.4
5.97
16.9
24.1
Box
0.92
1.38
3.04
14.6
15.0
19.3
8.85
7.92
12.6
Smooth-mask box
1.08
1.74
3.56
12.3
12.7
15.6
6.77
8.09
8.14
Box FT
0.81
1.13
2.37
15.2
15.6
18.4
8.70
9.07
11.6
Mixed-PGD FT
0.76
0.92
1.34
15.1
15.6
15.9
8.76
8.97
10.8
Table S16: Per-field LDC errors ( det , static draw 0, held-out samples, relative L2 error in %, means over five seeds; Digital: three seeds, ideal hardware). nom.: nominal device; MC: the MC device with the largest joint error among 200, that is, the device that sets the maximum in R200 ; search: the searched device. The joint error, the LDC metric, equals the pressure error to 0.001 percentage points in every model and device. The turbulent kinetic energy (TKE) k and the speed ∣u∣ are predicted far less accurately by all models, including the digital FNO. The searched device, selected on the joint error, raised the k error above that of the MC device with the largest joint error in 26 of 30 models and the ∣u∣ error in 24.
shared at zsearch
R200 , other boxes
Task
Arm (s0)
nominal
max 200
max 1000
searched
R200
R1000
crosstalk
extinction
32-D
8-D
method
Darcy
Random-error-aware
1.58
8.67
8.67
10.0
1.16
1.16
0.6
33
1.38
1.15
A
Box
2.46
3.29
3.60
5.77
1.75
1.60
0.4
27
1.99
1.37
Pc
Smooth-mask box
2.85
3.95
4.25
6.68
1.69
1.57
0.4
27
1.92
1.40
Pc
Box FT
2.27
3.16
3.27
5.19
1.64
1.59
0.4
27
1.95
1.35
Pc
Mixed-PGD FT
2.00
2.73
2.81
4.06
1.49
1.44
0.4
27
1.73
1.31
Pc
Table S17: Correlated tolerance box (seed 0, mode-level device, det , static draw 0, held-out samples, error in %). Crosstalk and extinction are shared by the four layers and the other six knobs are independent per layer ( D=26 ; Supplementary Note 9). MC: 1000 devices, maximum over the first 200 and over all 1000. Shared at zsearch : crosstalk width (pixel) and extinction ratio (dB) of the searched device. R200 , other boxes: the same model in the independent 32-D box (Supplementary Tables S1 and S2 ) and in the shared 8-D box (Supplementary Table S8 ; not run for HX). Method codes as in Supplementary Table S1 .
Qp
Qs
Qmax
Arm (seeds)
nom.
MC max
search
ratio
n>
nom.
MC max
search
ratio
n>
nom.
MC max
search
ratio
n>
Main round, mode-level device
Random-error-aware (5)
0.125
5.37
6.54
1.18–1.24
5/5
0.0391
2.95
3.61
1.21–1.24
5/5
0.160
3.06
3.66
1.15–1.26
5/5
Box (5)
0.132
0.367
0.539
0.46–2.21
3/5
0.0597
0.212
0.212
0.30–1.75
1/5
0.210
0.414
0.537
0.69–1.90
4/5
Smooth-mask box (5)
0.371
0.734
1.28
1.53–2.01
5/5
0.146
0.354
0.535
1.28–2.04
5/5
0.283
0.583
0.693
0.81–1.58
3/5
Box FT (5)
0.121
0.407
0.506
0.28–2.02
3/5
0.0516
0.228
0.180
0.27–1.51
1/5
0.222
0.403
0.486
0.52–1.84
3/5
Table S18: Heat exchanger: engineering quantities ( det , static draw 0, held-out samples). Qp : slice-mean gauge pressure; Qs : root-mean-square in-plane speed; Qmax : 99th percentile of the in-plane speed (Supplementary Note 9). Errors are ∣Qpred−Qtrue∣ in % of the test-set mean of ∣Qtrue∣ , averaged over the held-out samples, and are means over seeds for the nominal device (nom.), the largest over all 1000 MC devices (MC max) and the searched device. Ratio: searched error divided by the MC maximum, range over seeds (All: median). n> : models in which the searched error exceeds the MC maximum. The search maximises the field error, not these quantities.
mean removed
>1
seed 0, R1000mr
Arm (seeds)
R200
R200mr
R1000mr
R200p,mr
R200mr
R1000mr
offset share
emr/e
searched
re-searched
Random-error-aware (5)
1.40
1.39
1.35
1.39
5/5
5/5
0.94
0.98
1.35
1.36 (Pk)
Box (5)
2.22
1.63
1.59
1.63
5/5
5/5
0.93
1.00
1.59
1.73 (Pm)
Smooth-mask box (5)
1.89
1.76
1.65
1.76
5/5
5/5
0.92
1.05
1.80
1.83 (Pa)
Box FT (5)
2.08
1.56
1.53
1.56
5/5
5/5
0.92
1.03
1.48
1.52 (Pz)
Mixed-PGD FT (5)
1.44
1.15
1.15
1.15
5/5
5/5
0.84
1.32
1.15
1.25 (Pc)
Table S19: LDC with the pressure offset removed ( det , static draw 0, held-out samples; Supplementary Note 10). The spatial mean of the pressure is removed from each predicted and true field separately. R200 : original metric (Supplementary Table S2 ). R200mr , R1000mr : searched error divided by the MC maximum of 200 and 1000 devices on the mean-removed joint error; R200p,mr : on the mean-removed pressure error alone. Values are medians over seeds; >1 : models with the ratio above 1. Offset share: share of the squared norm of the searched device’s pressure error that lies in its spatial mean. emr/e : mean-removed over original joint error of the searched device. Seed 0: R1000mr of zsearch , found on the original metric, and of a PGD re-search on the mean-removed metric (start codes as in Supplementary Table S1 ; Pz: from zsearch ). Range: minimum to maximum over the 30 models.
Qp
Qs
Arm (seeds)
R1000
n>MC
search /W1
n>W1
Rfield
1/2/5%
R1000
n>MC
search /W1
n>W1
Rfield
1/2/5%
Main round, mode-level device
Random-error-aware (5)
1.24
5/5
1.40
5/5
1.01
0/0/5
1.24
5/5
1.40
5/5
1.01
0/0/0
Box (5)
2.45
5/5
3.33
5/5
1.43
1/0/0
2.34
5/5
3.29
5/5
2.66
0/0/0
Smooth-mask box (5)
1.91
5/5
2.99
5/5
1.09
5/0/0
1.83
5/5
3.07
5/5
1.32
0/0/0
Box FT (5)
2.34
5/5
3.32
5/5
1.75
1/0/0
2.27
5/5
3.54
5/5
3.00
0/0/0
Table S20: Heat exchanger: quantity-targeted search and Wilks 95/95 limits ( det , static draw 0, held-out samples; Supplementary Note 10). The search maximises the development error of Qp or Qs (7000 evaluations each). R1000 : searched error divided by the largest over all 1000 MC devices; search /W1 : divided by the first-order Wilks limit, the largest of the first 59 MC devices; Rfield : divided by the error of the same quantity on the field-searched device zsearch . Values are medians over seeds. n>MC : models whose searched error exceeds the MC maximum; n>W1 : models whose searched error exceeds W1 of all 16 disjoint sets of 59 MC devices. 1/2/5%: models in which W1 is at or below a threshold of 1, 2 or 5% of the mean quantity while the searched error exceeds it.
Figure S1: All 120 main models ( det , held-out samples): nominal error, MC mean to MC maximum of 1000 (line), MC maximum of 200 (bar), best unseen-attack device (open) and searched device (filled).
Figure S2: Exceedance curves of the 1000 MC devices. Fraction of the 1000 MC devices whose error exceeds c times the nominal error, one line per model ( det , held-out samples), with pointwise two-sided 95% Clopper–Pearson intervals. Open circles mark the upper end of the interval (‘95% upper bound’ in the legend) where no device exceeds the threshold. Red triangles mark c = searched error / nominal error. They sit at that upper end (0.37%, dotted line; main-text Methods) because no MC device reached the searched error in any model. Grey lines at the top are Random-error-aware models.
Figure S3: Error found by each search method relative to the best across methods. Best held-out error of the method divided by the best over all methods, mean over seeds, per task and arm (rows), including the circular-phase set (blank: not run).
Figure S4: Budget-matched random sampling. a–d , Best development error found so far by each sampler (corners + uniform: the mixed sampler), divided by the searched device’s development error, against the number of devices (seed-0 models, mode-level device). Lines are means over the six arms of a task, and bands the range over arms for the uniform and corner samplers. The dotted line is the uniform sampler of the Box model alone, continued to 100,000 devices, and the dashed vertical line marks the search budget. e , Rmatch of each sampler for the 24 models, six arms per task in the order of main-text Fig. 4. f , Rmatch of the uniform, corner and mixed samplers on the pixel-level device (R, Random-error-aware; B, Box; M, Mixed-PGD FT; three seeds each) and for the four transferred models (T). Dashed line: Rmatch=1.1 .
Figure S5: HX: joint error against the primary metric. Joint relative L2 error (median over held-out samples) against the primary metric (mean of the per-channel relative L2 errors) for every HX model on the nominal (faint) and searched (solid) device, and for the digital FNO on ideal hardware. Both axes are logarithmic, and the dashed line marks equality.
Figure S6: LDC field snapshot (mode-level device, det , held-out sample with the median nominal error of the Box seed-0 model). a , Box seed 0. b , Mixed-PGD FT seed 0. True fields and errors (prediction minus true field) on the nominal device, on the worst of the first 200 MC devices (selected by development error) and on the searched device; the lid is at the top. Rows: pressure p , turbulent kinetic energy k and velocity magnitude ∣u∣ (Supplementary Note 3). The pressure dominates the joint error used as the LDC metric (joint errors, equal to the pressure errors shown: 0.40%, 0.94% and 2.39% for Box; 0.35%, 0.47% and 0.94% for Mixed-PGD FT). Error maps of a field share a symmetric colour scale, clipped at the 99.5th percentile of the absolute error, and give that field’s relative L2 error. True fields span their 0.5th to 99.5th percentiles. Colour-bar values are to be multiplied by the power of ten in the row label.
Figure S7: Further field snapshots , as in main-text Fig. 5b. a , HX Random-error-aware seed 0 (mode-level device; mean of the per-channel errors 0.22%, 3.61% and 4.91%). Its errors on the MC and searched devices are smooth and large-scale, unlike the fine-grained errors of the Box model in main-text Fig. 5b. b , Darcy Box seed 0 on the pixel-level device (cell integral): the error grows in the same region on all three devices (1.88%, 2.09% and 2.40%). Colour scales as in Supplementary Fig. S6 .
Figure S8: Ratio of the local linear estimate to the searched error (full test set, det ).
Figure S9: Architectures of the two tasks with vector inputs. Tensor shapes are channels × height × width on grids and nodes × features on the mesh. Blue: wave-based (4f optical) paths; amber: digital. a , LDC. The 90-dimensional encoding of the lid-velocity history passes through a digital MLP ( 90→256→256→8×9×9 , GELU after the first two maps), is reshaped to an 8×9×9 latent field and is upsampled bilinearly to 8×65×65 . The hybrid FNO of main-text Fig. 1a appends two grid coordinates, lifts to 24 channels and applies the four hybrid layers. The projection ( 24→128→3 ) gives (p,k,∣u∣) on the 65×65 grid. b , HX. The 102-dimensional condition vector passes through the same encoder onto a 65×65 latent grid over [−1,1]2 . The 24 features of the last hybrid layer are sampled bilinearly at the 3977 node coordinates and concatenated with a 26-dimensional coordinate encoding. A per-node digital MLP ( 50→128→3 ) replaces the projection and gives the pressure and the two in-plane velocities (p,u1,u2) at the nodes. The mesh inset is a schematic, not the real node coordinates.
Figure S10: Certificate bound divided by the highest error found , against the box scale κ (v2 simulator, Darcy seed 0). Line style: 32 knobs or extinction fixed. Markers: exact arithmetic or float32 mask slack. Dotted line: 3 × .
Figure S11: v1 pilot: per-pixel static errors (Darcy) for the v1 Random-error-aware and Box seed-0 models. Error without static errors, mean over 20 random draws inside the per-pixel box, and the error after gradient search over all 2.88 million per-pixel variables (legacy simulator).
Figure S12: Cross-chip penalty when adapters calibrated on one simulated chip are used on others , for a learned frozen library and for random frozen optics (frozen-library pilot, v1 simulator; Ntrain : number of training samples).
Figure S13: Heat exchanger: errors of three engineering quantities for all 39 HX models ( det , static draw 0, held-out samples; Supplementary Note 9). Error of the quantity in % of its test-set mean on the nominal device, from the mean to the maximum over all 1000 MC devices, and on the searched device, which maximises the field error. Filled diamonds: the searched error exceeds the MC maximum; open diamonds: it does not. Colour gives the arm. The y -axes are logarithmic.
Figure S14: Heat exchanger: quantity-targeted search against Wilks 95/95 limits for all 39 HX models ( det , static draw 0, held-out samples; Supplementary Note 10). Error of the quantity in % of its test-set mean: the first-order Wilks limit W1 from the first 59 MC devices (open circles) and its range over the 16 disjoint sets of 59 (bars), the maximum over all 1000 MC devices (dashes) and the device found by the search for that quantity (diamonds; colour gives the arm). Dotted lines: thresholds of 1, 2 and 5%. The y -axes are logarithmic.
Machine-learning surrogates accelerate physical simulation, but lower prediction error need not coincide with lower error in physically relevant flow statistics. We examine this question for flow around a NACA4418 airfoil using paired computational-fluid-dynamics simulations and experimental particle-image-velocimetry measurements. A mean-preserving input intervention removes velocity fluctuations from selected regions of observed flow histories. Across four neural operators, removing fluctuations from the most energetic 10% of valid observed cells changes forecasts more than equal-area random removal. Because the masks are not matched for removed fluctuation energy, this contrast measures sensitivity, not independent evidence of physical importance. Separately, a CNO has lower velocity-field error but substantially higher two-component fluctuation-energy error than the reference on both analysis subsets. An output attenuation stress test also demonstrates disagreement between benchmark errors and domain-summed fluctuation energy. These single-benchmark results motivate reporting complementary physical diagnostics alongside aggregate prediction scores; they do not establish counterfactual physical correctness.
Somyajit Chakraborty, Xizhong Chen
State Key Laboratory of Synergistic Chem-Bio Synthesis, Department of Chemical Engineering, School of Chemistry and Chemical Engineering, Shanghai Jiao Tong University, Shanghai 200240, People’s Republic of China
Neural operators are data-driven models that learn mappings from inputs that parameterize partial differential equations, such as spatially varying coefficients, initial conditions, forcing terms, boundary conditions, or geometries, to solution fields or quantities of interest. Once trained, they can serve as surrogates for classical numerical solvers in many-query settings that require repeated evaluations for varying inputs. We address the question of when, and then why, neural operator surrogates outperform classical numerical solvers, in terms of cost for a given accuracy. We focus on the post-training, many-query limit, in which data-acquisition and training costs are treated as fixed and fully amortized. Even in this deliberately favorable regime for neural operators, there are regimes in which classical solvers outperform the surrogate models. We compare the cost-accuracy performance of neural operator surrogates and classical numerical solvers through a reproducible benchmark study comparing neural operators with problem-matched classical solvers on representative problems in computational science and engineering, focusing on prediction error, per-query floating-point cost, and wall-clock runtime. Neural operators are most competitive at low-to-moderate accuracy requirements. Their floating-point cost advantage depends strongly on the problem structure, arising when they avoid temporal or nonlinear iterations or predict a reduced quantity of interest rather than a full solution field. Additional wall-clock speedups result from dense tensor operations that are well suited to modern hardware. As the target accuracy is tightened, achieving the required accuracy with neural operators becomes increasingly challenging, and classical solvers outperform surrogates in this regime; thus classical solvers will remain important for verification and high-accuracy computation.
Daniel Zhengyu Huang, Andrew M. Stuart
Beijing International Center for Mathematical Research, Center for Machine Learning Research, Peking University, Beijing, China · California Institute of Technology, Pasadena, CA
Physics-informed neural networks and hybrid models infer PDE coefficients from noisy data. When a trained network returns one, no standard check says whether to trust it. We show what those checks report when the operator is wrong: one sensor aggregating several diffusion sources. On one parabolic benchmark at 2% noise, the in-domain error is 1.4 times the noise while the identified diffusivity settles 30% off. Every least-squares minimiser reaches that value, which drifts 27% across windows; the network, whose objective is composite, settles 1.3% away. The checks stay as silent when the design is blind to a rate of a richer operator, though the remedies are opposite. We develop a reference-free diagnostic, read in the physical parameter, not the weights, without retraining the network: an information-matrix test on the residuals, a heterogeneity statistic across window refits, and a Fisher-rank statistic on the design at the rates the single fit postulates. On the analytic head the specification test holds its pre-registered ceiling and rejects every misspecified replicate of both benchmark configurations, with a notch against a missing reaction term. The rank statistic is exactly zero only where the design is blind; a wrong operator confined to that mode leaves the specification test mute, and the rank statistic says so before any fit. The window reading exceeds its ceiling by one seed in thirty. A network frozen at its minimum returns the same verdicts; one stopped short rejects as a wrong operator would.
Eric Fock
PIMENT Laboratory, Universit´e de La R´eunion, Le Tampon 97430, La R´eunion, France