Neural surrogates have emerged as fast alternatives to the numerical simulation of three-dimensional turbulence. However, training them at high resolution remains challenging, since the memory of full-field models grows with the resolution. In addition, full-resolution training data are expensive to simulate and store, and therefore scarce. We introduce ScaleSplit-NO (Scale-Split Neural Operator), which exploits the scale structure of turbulence with two neural operators: a Parent predicts the global coarse field at the next time step, and a Child predicts full-resolution local patches conditioned on this prediction. Neither model operates on the full-resolution field. The Child is pretrained alone and then attached to the Parent's coarse prediction through zero-initialized connections. On two complex high-resolution turbulence benchmarks, ScaleSplit-NO surpasses all competing baselines in both prediction accuracy and data efficiency. On the higher-resolution dataset JHTDB256 (2563), its normalized mean squared error (NMSE) is 53% lower than that of the strongest baseline, and its training memory is 79% lower than that of the most memory-efficient baseline. We further demonstrate its effectiveness for urban wind prediction in a real district of Montreal on a 500×150×500 grid, reducing one-step NMSE by 65.8% relative to the baseline. Moreover, swapping in a Parent trained on additional coarse fields improves prediction without retraining the Child, providing further accuracy gains at a small storage cost.
Figures & tables
Figure 1: Accuracy versus training memory. Test error (NMSE) with limited training data against the peak GPU memory during training (Table 1 ). Lower left is better.
Figure 2: Separate supervision and composed prediction. (a) The Parent is trained to predict the coarse target Dut+1 from the coarse input Dut . (b) The Child is pretrained to predict the target patch Tjut+1 from the input patch Tjut . (c) The coarse prediction of the frozen Parent is interpolated and cropped to the condition cj , which enters the Child through zero-initialized connections (diamonds). Fine-tuning updates the backbone and these connections, and the loss ℓ is the only place where the target field enters. At inference all weights are fixed, and the predicted patches y^j are assembled into the predicted field u^t+1=F(ut) .
NMSE, low data (1/8)
NMSE, full data
Training
One step
Rollout
One step
Rollout
Memory
Method
×10−3
×10−2
×10−3
×10−2
GiB
MHD64
U-Net
15.5
2.83
5.97
1.30
4.6
FNO
97.6
( ↑ 529%)
17.2
( ↑ 507%)
29.5
( ↑ 394%)
4.96
( ↑ 283%)
5.5
( ↑ 20%)
MG-TFNO
10.3
( ↓ 33%)
2.06
( ↓ 27%)
5.98
(0%)
1.27
( ↓ 2%)
10.6
( ↑ 130%)
EddyFormer
8.83
( ↓ 43%)
1.99
( ↓ 30%)
5.88
( ↓ 2%)
1.32
( ↑ 2%)
17.0
( ↑ 268%)
Table 1: Test error (NMSE) and training memory. U-Net serves as the reference. Parentheses give the change in error and memory relative to U-Net, ( ↓ better) or ( ↑ worse) . Bold : best result. Thick underline : best baseline. Dashed underline : second-best baseline.
Figure 3: Visualization of 3D turbulence predictions and errors. The figure shows one test pair of each dataset for one-step prediction in the low-data setting. Each panel draws a 3D field as surfaces of constant value (isosurfaces) at the levels in its legend. Velocity panels show the x -velocity of the ground truth and of the prediction of ScaleSplit-NO. Error panels show the magnitude of the velocity error ∥u^t+1−ut+1∥ of ScaleSplit-NO and of the strongest baseline of each dataset.
Figure 4: Test error versus the amount of stored training data. (a1, a2) Test error of ScaleSplit-NO and the two strongest baselines of each dataset for different amounts of training data. (b1, b2) Test error of ScaleSplit-NO when a Child trained on the labeled fraction of the data is kept fixed and combined with Parents trained on additional coarse fields. In (b1, b2), gray dashed lines connect models whose Parent and Child are trained on the same data, and the relative training data size also counts the additional coarse fields.
Figure 6Figure 7
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
MHD64
JHTDB256
Method
Low data (1/8)
Full data
Low data (1/8)
Full data
U-Net
11.3
7.01
9.94
8.14
FNO
29.0
15.9
30.3
19.8
MG-TFNO
9.37
7.06
16.4
15.3
EddyFormer
8.58
6.94
19.5
18.8
P3D
11.2
7.10
5.07
4.62
Appendix
Table 4: One-step test NRMSE ( ×10−2 ) with 1/8 and with all of the training data.
Figure 7: MHD64, x-velocity of all methods. One-step prediction (left) and rollout step 2 (right) with the normalized errors.
Figure 8: MHD64, all seven channels at one step. Ground truth and normalized errors of the two strongest baselines, MG-TFNO and EddyFormer, and of ScaleSplit-NO.
Figure 9: MHD64, complete field at one step. Isosurfaces of the x-velocity at one and two standard deviations of the ground truth, and of the velocity error magnitude at 0.10 and 0.16, for MG-TFNO, EddyFormer, and ScaleSplit-NO.
Figure 10: JHTDB256, x-velocity of U-Net, FNO, MG-TFNO, and ScaleSplit-NO. One-step prediction (left) and rollout step 15 (right) with the normalized errors.
Figure 11: JHTDB256, x-velocity of EddyFormer, P3D, ReViT, and ScaleSplit-NO. One-step prediction (left) and rollout step 15 (right) with the normalized errors, on the same scales as Figure 10 .
Figure 12: JHTDB256, all four channels at one step. Ground truth and normalized errors of the two strongest baselines, U-Net and P3D, and of ScaleSplit-NO.
Figure 13: JHTDB256, complete field at one step. Isosurfaces of the x-velocity at one and two standard deviations of the ground truth, and of the velocity error magnitude at 0.15 and 0.30, for U-Net, P3D, and ScaleSplit-NO.
Figure 14: Urban wind at one step. Isosurfaces of the fluctuation of the velocity component v about its temporal mean at ±1 and ±2 , and of the absolute error of v at 0.3 and 0.6, for U-Net and ScaleSplit-NO. The NMSE refers to v in the fluid region of this test pair.
Figure 15: MHD64, velocity spectra at one step. Energy spectra of the ground truth and of all predictions (left) and of the prediction errors (right).
Figure 16: JHTDB256, velocity spectra at one step. Energy spectra of the ground truth and of all predictions (left) and of the prediction errors (right).
Method
MHD64
JHTDB256
U-Net
21.45±0.34
17.19±0.86
FNO
99.81±2.23
99.50±10.61
MG-TFNO
9.67±0.34
32.23±3.91
EddyFormer
12.45±0.23
41.10±0.02
P3D
21.15±0.94
4.07±0.04
ReViT
93.89±0.13
32.96±0.09
Appendix
Table 5: One-step test NMSE ( ×10−3 ) of all methods with 1/8 of the data over three repeated runs at 20% of the training cost (seeds 101, 202, 303), mean ± sample standard deviation.
NMSE
ScaleSplit-NO final configuration
6.88±0.11
w/o pretraining and zero init.
7.22±0.23
w/o Child fine-tuning
7.23±0.04
Fine-tuning with coarse target
12.04±0.48
Residual correction
7.33±0.14
Joint Parent–Child training
11.81±0.02
Appendix
Table 6: One-step test NMSE ( ×10−3 ) of the ablation variants of Table 5 on MHD64 with 1/8 of the data over three repeated runs at 20% of the training cost, mean ± sample standard deviation.
Parent
Dataset
Child
1/8
1/4
1/2
1
Coarse target
MHD64
1/8
6.20
5.68
5.10
4.99
2.63
1/4
5.53
4.94
4.82
2.44
1/2
4.79
4.77
2.48
1
4.52
2.35
JHTDB256
1/8
1.254
1.248
1.238
1.231
1.182
Appendix
Table 7: One-step test NMSE ( ×10−3 ) of a fixed Child with Parents trained on more data, and with the coarse target as condition. Rows: training data of the Child. Columns: training data of the Parent. Both are given as fractions of the full training set.
Dataset
Setting
NMSE
MHD64
Final: global coarse grid = Parent grid 243 , local Child grid 163
6.20
Global coarse grid = Parent grid 243→163
6.20
( ↑ 0.1%)
Global coarse grid = Parent grid 243→323
6.42
( ↑ 3.5%)
Local Child grid 163→123
6.30
( ↑ 1.6%)
Local Child grid 163→243
7.06
( ↑ 13.9%)
JHTDB256
Final: global coarse grid 643 , Parent grid 323 , local Child grid 163
1.25
Appendix
Table 8: Sizes of the Parent and the Child with 1/8 of the data. One-step test NMSE ( ×10−3 ), with the change relative to the final configuration of each dataset in parentheses, ( ↑ worse) .
Overlap
NMSE ( ×10−3 )
Voxels
Time (s)
0%
1.29
1.0 ×
3.44
33%
1.25
3.4 ×
3.46
50%
1.25
8.0 ×
3.50
Appendix
Table 9: Overlap of the Parent patches on JHTDB256 with 1/8 of the data. Bold : the default configuration.
Parent
Child
MHD64
JHTDB256
MHD64
JHTDB256
Input
243 field
323 patch
163 patch
163 patch
Fourier blocks
4
4
8
8
Hidden width
32
32
64
64
Modes per axis
12
16
8
8
Parameters ( 106 )
56.6
134.2
134.3
134.3
Appendix
Table 10: Backbone configurations. Modes are the number of Fourier modes per axis. The parameters of the Child include the condition pathway.
Dataset
Stage
Epochs
Learning rate
Batch
MHD64
Parent training
300
2×10−3
1 field
Child pretraining
50
1×10−3
8 patches
Child fine-tuning
300
2×10−3
8 patches
JHTDB256
Parent training
300
2×10−3
8 patches
Child pretraining
100
5×10−4
16 patches
Child fine-tuning
500
2×10−3
16 patches
Appendix
Table 11: Training settings of the three stages.
Method
Width
Depth
Modes
Epochs
Learning rate
Batch
U-Net
64 / 16
4 levels
–
300
2×10−4
2 / 1
FNO
48 / 12
4 layers
16 / 12
400
1×10−3
1
MG-TFNO
64 / 40
4 layers
32 / 24
500 / 200
1×10−3
1 field
EddyFormer
32
4 layers
13
10k–80k steps
1×10−3
1
P3D
P3D-L / P3D-B
–
1,000 / 4,000
2×10−4
8 / 4
ReViT
48
1-2-4-2-1
–
300
1×10−3
2 / 1
Appendix
Table 12: Baseline configurations on MHD64 / JHTDB256. Width is the base channel count of U-Net, the hidden width of FNO, MG-TFNO, and EddyFormer, and the embedding dimension of ReViT. EddyFormer is trained for 10k steps with 1/8 of the data and 80k steps with all of the data.
MHD64
JHTDB256
Input
Batch
Memory
Time
Input
Batch
Memory
Time
Method
GiB
h
GiB
h
U-Net
643
2
4.6
4.6
2563
1
31.7
41.3
FNO
643
1
5.5
4.6
2563
1
35.0
41.4
MG-TFNO
323
8
10.6
20.1
1283
1
19.8
32.7
EddyFormer
643
1
17.0
16.7
2563
1
26.4
22.9
Appendix
Table 13: Training samples. Input is the grid size of one training sample, batch the number of samples per update, memory the peak GPU memory allocated during one training update, and time the training time with all of the data. For ScaleSplit-NO the first row gives the maximum memory and the total time of its three stages, which are trained one after another and listed below it. MG-TFNO splits each field into eight patches and computes one loss on the prediction assembled from all of them, so every update uses one whole field. For MG-TFNO, Input is the size of one patch without padding, 483 and 1443 with periodic padding, and Batch the number of patches processed together.
MHD64
JHTDB256
Method
Overlap
Time (ms)
Overlap
Time (s)
U-Net
–
28
–
0.31
FNO
–
15
–
0.41
MG-TFNO
–
55
–
0.71
EddyFormer
–
341
–
0.38
P3D
–
35
–
0.41
Appendix
Table 14: Inference time. Time for one complete field. For ScaleSplit-NO the time is given for several overlaps of the Child patches.
Figure 17: Training memory versus batch size. Peak GPU memory of one training update, with the input size of one sample in the legend. ScaleSplit-NO is shown for each of its three training stages.
Evaluating neural operators for 3D turbulent flow requires validated datasets with physical benchmarks. We present a reproducible pipeline generating training data for 3D channel flows around generated geometries at Re=1,000-10,000. Our lattice Boltzmann solver with cumulant collision operators is rigorously verified against experimental measurements (Strouhal number, drag coefficients, turbulent fluctuations) with comprehensive grid convergence studies at resolution 1024x512x512. Building upon an established framework, this validated pipeline enables standardized surrogate model comparison. We outline planned systematic evaluation of Fourier Neural Operator and U-Net variants on forecasting, super-resolution, and error correction tasks, using physics-informed metrics to assess turbulent energy cascade representation. Future work will compare computational efficiency between numerical solvers and neural surrogates, exploring practical application. We seek community feedback on our validation approach, planned benchmark methodology, and evaluation priorities for neural operators in turbulent flows.
Lukas Schröder, Shubham Kavane, Harald Köstler
Chair of Computer Science 10 (System Simulation) Friedrich-Alexander-Universität Erlangen-Nürnberg
In this paper, we propose a perturbation-based conformal prediction framework for uncertainty quantification in operator learning, with a focus on the 2D Navier--Stokes equations. While neural operators provide fast surrogates for expensive PDE solvers, they do not by themselves provide calibrated uncertainty for spatiotemporal field predictions. Our approach wraps a trained Fourier Neural Operator (FNO) with split conformal prediction and constructs the local uncertainty scale by comparing the predictions of two operators trained on nearly identical datasets: one on the original labels and one on labels perturbed by small Gaussian noise. We consider this procedure in the data-scarce regime, where the total label budget is fixed and methods that require a separate uncertainty network must divide training data between multiple models. On the 2D Navier--Stokes benchmark, the perturbation-based method produces substantially narrower conformal bands than existing methods under matched total data budgets while maintaining the target simultaneous coverage. These results suggest that perturbation sensitivity is a practical and sample-efficient uncertainty proxy for conformalized neural operators.
Weinan Wang, Bowen Gang, Hao Deng
Department of Mathematics, University of Oklahoma, Norman, OK, USA · Department of Statistics and Data Science, School of Management, Fudan University, Shanghai, China
Neural operators serve as fast, data-driven surrogates for scientific modeling but typically rely on a monolithic, single-pass inference procedure that struggles to resolve high-frequency details, a limitation known as spectral bias. We introduce the Iterative Refinement Neural Operator (IRNO), which augments pre-trained operators with a learned refinement module iteratively applied via fixed-point iteration. IRNO decomposes the prediction into a coarse initialization followed by successive residual corrections, paralleling classical numerical solvers. Under local assumptions, we establish contraction of the induced operator, ensuring convergence to a unique fixed point. To explicitly target high-frequency errors, we propose a progressive spectral loss that adaptively increases penalty on high-frequency components over refinement steps during training. Across physical systems, IRNO consistently lowers error, with up to 56.05% improvement on turbulent flow. On Active Matter, spectral analysis reveals that, relative to base operator, the normalized error ratios decrease to 27.72-36.10% in low-, 5.07-6.68% in mid-, and 1.48-2.04% in high-frequencies, remaining stable beyond the trained iteration count. Code is available at https://github.com/xiaotianliu-dartmouth/Iterative_Refinement_Neural_Operator
Xiaotian Liu, Shuyuan Shang, Xiaopeng Wang +2
1Dartmouth College · 2CUHK Shenzhen · 3Lawrence Berkeley National Lab