cs.LGOct 5, 2026

Targeted search shows that random-device testing underestimates worst-case error in a simulated wave-based neural operator

Authors: Samrendra Roy, Jason Yoo, Souvik Chakraborty, Syed Bahauddin Alam

Organizations: Department of Nuclear, Plasma, and Radiological Engineering, University of Illinois Urbana-Champaign, Urbana, IL 61801, USA · Department of Applied Mechanics, Indian Institute of Technology Delhi, Hauz Khas, New Delhi 110016, India · Yardi School of Artificial Intelligence (ScAI), Indian Institute of Technology Delhi, Hauz Khas, New Delhi 110016, India · National Center for Supercomputing Applications (NCSA), University of Illinois Urbana-Champaign, Urbana, IL 61801, USA

Abstract

Wave-based processors promise fast, energy-efficient Fourier layers for neural operators. They are usually validated on randomly sampled devices, but using them requires knowing how large their error can become under fabrication and alignment variation. In a stylised numerical case study, a hybrid Fourier neural operator runs its four spectral layers on simulated coherent 4f processors with 32 toleranced knobs, whose half-widths are representative rather than calibrated. For 120 models (four tasks, six training methods, five seeds), we compared the worst of N random in-spec devices with a searched one. On a deterministic simulator with one frozen draw of the random static errors, the searched device's held-out error was 1.08-3.10 times the maximum over 200 Monte Carlo devices and 1.06-2.71 times that over 1000. With 20 fresh static draws, it still exceeded the maximum over 200 random devices in 116 of 120 models. Under uniform sampling, the probability of drawing such a device is at most 0.37% per model (two-sided 95% Clopper-Pearson), which says nothing about how large its error is. The gap persisted with uniform or Sobol' sampling at the search's budget, shared knobs, a second crosstalk model, box scales of 0.25-2 and a pixel-level device model. Models trained only with random static errors reached 3.7-39.9 times their nominal error on searched devices, and fine-tuning on random and gradient-searched devices gave the lowest searched error of the six in all 20 task-seed pairs. For two heat-exchanger quantities, a search targeted at each exceeded the worst of 1000 random devices in all 39 models, and hence the Wilks 95/95 limit (worst of 59). For the mean pressure of 11 models, no random device exceeded a 1% error threshold, but the searched device did. Random testing estimates how often errors exceed a threshold; worst-device search gives a lower bound on how large they can be.

Figures & tables

Explore similar work

Sep 30, 2026cs.LG

Do Better Scores Mean Better Physics? Physics-Grounded Explanations for Sim2Real Neural Operators

Machine-learning surrogates accelerate physical simulation, but lower prediction error need not coincide with lower error in physically relevant flow statistics. We examine this question for flow around a NACA4418 airfoil using paired computational-fluid-dynamics simulations and experimental particle-image-velocimetry measurements. A mean-preserving input intervention removes velocity fluctuations from selected regions of observed flow histories. Across four neural operators, removing fluctuations from the most energetic 10% of valid observed cells changes forecasts more than equal-area random removal. Because the masks are not matched for removed fluctuation energy, this contrast measures sensitivity, not independent evidence of physical importance. Separately, a CNO has lower velocity-field error but substantially higher two-component fluctuation-energy error than the reference on both analysis subsets. An output attenuation stress test also demonstrates disagreement between benchmark errors and domain-summed fluctuation energy. These single-benchmark results motivate reporting complementary physical diagnostics alongside aggregate prediction scores; they do not establish counterfactual physical correctness.
Sep 21, 2026math.NA

Cost-Accuracy Trade-offs: Neural Operator vs Classical Numerical Solver

Neural operators are data-driven models that learn mappings from inputs that parameterize partial differential equations, such as spatially varying coefficients, initial conditions, forcing terms, boundary conditions, or geometries, to solution fields or quantities of interest. Once trained, they can serve as surrogates for classical numerical solvers in many-query settings that require repeated evaluations for varying inputs. We address the question of when, and then why, neural operator surrogates outperform classical numerical solvers, in terms of cost for a given accuracy. We focus on the post-training, many-query limit, in which data-acquisition and training costs are treated as fixed and fully amortized. Even in this deliberately favorable regime for neural operators, there are regimes in which classical solvers outperform the surrogate models. We compare the cost-accuracy performance of neural operator surrogates and classical numerical solvers through a reproducible benchmark study comparing neural operators with problem-matched classical solvers on representative problems in computational science and engineering, focusing on prediction error, per-query floating-point cost, and wall-clock runtime. Neural operators are most competitive at low-to-moderate accuracy requirements. Their floating-point cost advantage depends strongly on the problem structure, arising when they avoid temporal or nonlinear iterations or predict a reduced quantity of interest rather than a full solution field. Additional wall-clock speedups result from dense tensor operations that are well suited to modern hardware. As the target accuracy is tightened, achieving the required accuracy with neural operators becomes increasingly challenging, and classical solvers outperform surrogates in this regime; thus classical solvers will remain important for verification and high-accuracy computation.
Aug 4, 2026cs.LG

Wrong Operator or Blind Design? A Reference-Free Diagnostic for Physics-Informed Coefficient Learning

Physics-informed neural networks and hybrid models infer PDE coefficients from noisy data. When a trained network returns one, no standard check says whether to trust it. We show what those checks report when the operator is wrong: one sensor aggregating several diffusion sources. On one parabolic benchmark at 2%2\% noise, the in-domain error is 1.41.4 times the noise while the identified diffusivity settles 30%30\% off. Every least-squares minimiser reaches that value, which drifts 27%27\% across windows; the network, whose objective is composite, settles 1.3%1.3\% away. The checks stay as silent when the design is blind to a rate of a richer operator, though the remedies are opposite. We develop a reference-free diagnostic, read in the physical parameter, not the weights, without retraining the network: an information-matrix test on the residuals, a heterogeneity statistic across window refits, and a Fisher-rank statistic on the design at the rates the single fit postulates. On the analytic head the specification test holds its pre-registered ceiling and rejects every misspecified replicate of both benchmark configurations, with a notch against a missing reaction term. The rank statistic is exactly zero only where the design is blind; a wrong operator confined to that mode leaves the specification test mute, and the rank statistic says so before any fit. The window reading exceeds its ceiling by one seed in thirty. A network frozen at its minimum returns the same verdicts; one stopped short rejects as a wrong operator would.