Simulation and experimental measurements provide complementary data for learning spatiotemporal physical systems, but standard simulation-to-experiment fine-tuning optimizes only the experimental objective after transfer and can degrade simulation performance. We formulate simulation--experiment prediction as a multi-objective learning problem with domain-specific simulation and experimental risks. On four fluid systems from RealPDEBench and two model capacities, we compare Simulation only, Experiment only, Sim→Exp, and Joint training, evaluating every final model on both held-out domains. Sim→Exp tends to specialize more strongly to experimental data at the cost of simulation-domain forgetting. Joint training consistently achieves the best balanced performance over a broad range of simulation--experiment evaluation weightings, while substantially improving simulation retention over Sim→Exp. Joint also better preserves simulation-only fields absent from experimental measurements. Project page: https://mahindrautela.github.io/morph.
Figures & tables
Dataset
Sim. fields
Exp. fields
Sim. grid
Exp. grid
Cylinder
u,v,p
u,v
64×128
128×256
Controlled Cylinder
u,v,p
u,v
64×128
128×256
FSI
u,v,p
u,v
128×128
128×128
Foil
u,v,p
u,v
128×256
128×256
Table 1: RealPDEBench systems. The four fluid datasets share velocity observables (u,v) ; simulation additionally contains pressure. The grids shown are the native benchmark grids.
Simulation only
Experiment only
Sim → Exp
Joint
Dataset
Capacity
bRMSE
bRelL2
bfRMSE
bRMSE
bRelL2
bfRMSE
bRMSE
bRelL2
bfRMSE
bRMSE
bRelL2
bfRMSE
Cylinder
17.5M
0.118098
0.539958
0.019925
0.101539
0.332502
0.015969
0.100507
0.419261
0.015714
0.058412
0.126032
0.009489
Cylinder
48.2M
0.096126
0.427966
0.015923
0.094315
0.311948
0.014704
0.094171
0.354185
0.014653
0.056230
0.098416
0.009077
Ctrl Cyl
17.5M
0.049899
0.333784
0.012611
0.061036
0.405474
0.014763
0.071465
0.490167
0.017985
0.030442
0.216437
0.006148
Ctrl Cyl
48.2M
0.048514
0.298634
0.012441
0.058852
0.393601
0.014363
0.077844
0.481965
0.019656
0.020455
0.149020
0.004180
FSI
17.5M
0.060337
0.405078
0.009692
0.082525
0.463016
0.013427
0.069920
0.392799
0.010924
0.033517
0.229200
0.004630
Table 2: Main result: balanced dual-domain performance on the four fluid datasets. bRMSE, bRelL2, and bfRMSE give simulation and experiment equal weight according to Eq. ( 1 ). Lower is better. Bold marks the best protocol for each dataset and model capacity. Joint is best in all 24 comparisons (8 settings × 3 metrics).
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Input frames
Target frames
Mask probability
Noise scale
Cylinder
20
20
0.1
0.1
Controlled Cylinder
10
10
0.1
0.1
FSI
20
20
0.5
0.1
Foil
20
20
0.1
0.1
Appendix
Table 3: Dataset-specific temporal and numerical-training settings used by the final runs. Masking and noise are applied only by the RealPDEBench numerical training dataset; validation, test, experimental training, and normalization data are clean.
Architecture
Variable grid
Variable field set
Field-wise cross-attention
Selected
Vanilla ViT [ Dosovitskiy et al., 2021 ]
Partial
No
No
No
U-Net [ Ronneberger et al., 2015 ]
Partial
No
No
No
FNO [ Li et al., 2020b ]
Yes
No
No
No
DPOT [ Hao et al., 2024 ]
Partial
Partial
No
No
Poseidon [ Herde et al., 2024 ]
Yes
Partial
No
No
Walrus [ McCabe et al., 2025 ]
Yes
Partial
No
No
Appendix
Table 4: Capability-based architecture selection for heterogeneous simulation–experiment training. “Partial” denotes that padding, interpolation, resizing, or task-specific interfaces are required in the standard formulation. The table compares interface capabilities rather than predictive accuracy.
Configuration
Parameters
Conv. filters
Latent dim.
Heads
Depth
MLP dim.
Ti
17,528,776
8
256
4
4
1024
S
48,161,736
8
512
8
4
2048
Appendix
Table 5: Architecture configurations used in the controlled study. Parameter counts are the trainable counts of the instantiated models.
Setting
Value
Optimizer
Adam
Learning rate
10−4
Weight decay
0
Scheduler
cosine annealing to zero, Tmax=10,000 per stage
Stage budget
10,000 optimizer updates
Physical train batch size
6 (Cylinder, Controlled Cylinder, FSI); 3 (Foil)
Appendix
Table 6: Optimization settings used in the final runs. The physical training batch size is dataset-specific but fixed across protocols within each dataset.
Protocol
Optimizer steps
Sim. batches
Exp. batches
Stage structure
Simulation only
10,000
10,000
0
one simulation stage
Experiment only
10,000
0
10,000
one experimental stage
Sim → Exp
20,000
10,000
10,000
10k sim + 10k exp
Joint
10,000
10,000
10,000
paired losses before each step
Appendix
Table 7: Optimizer-step and source-exposure budgets. Within each dataset, Joint and Sim → Exp are matched in source-specific mini-batch and sample exposure and in forward/backward evaluations, but not in optimizer-step count.
Simulation test
Experimental test
Dataset
Model
Params.
Training
RMSE ↓
Rel. L2↓
fRMSE ↓
RMSE ↓
Rel. L2↓
fRMSE ↓
Cylinder
MORPH-Ti
17.5M
Simulation only
0.014863
0.089989
0.002119
0.166353
0.989927
0.028099
Experiment only
0.108947
0.551503
0.017663
0.093546
0.113501
0.014074
Sim → Exp
0.105315
0.730744
0.017025
0.095457
0.107778
0.014283
Joint
0.017445
0.107683
0.002650
0.080744
0.144381
0.013156
MORPH-S
48.2M
Simulation only
0.015774
0.085703
0.002150
0.135024
0.770230
0.022416
Appendix
Table 8: Full dual-domain test performance on the four fluid datasets. Both simulation and experimental metrics are computed on the shared velocity fields (u,v) . Lower is better. Bold marks the best protocol for each model capacity, test domain, and metric. Results are seed 0.
Figure 1: Cylinder: sensitivity to simulation–experiment evaluation weighting. Rows show MORPH-Ti and MORPH-S and columns show RMSE, relative L2 , and fRMSE (lower is better). Here, λ=0 weights only experimental performance and λ=1 only simulation performance; λ=0.5 recovers the balanced metrics in Table 2 . Joint gives the lowest error for all three metrics and both capacities throughout λ=0.1 – 0.9 . At the single-domain endpoints, Sim → Exp is lowest only for experimental relative L2 , while Simulation only becomes best for MORPH-Ti at λ=1 ; Joint remains best for MORPH-S.
Figure 2: Controlled Cylinder: sensitivity to simulation–experiment evaluation weighting. Rows show MORPH-Ti and MORPH-S and columns show RMSE, relative L2 , and fRMSE (lower is better). Joint gives the lowest error for all three metrics and both capacities throughout λ=0.1 – 0.9 . Sim → Exp is best at the purely experimental endpoint ( λ=0 ), whereas at λ=1 Simulation only gives the lowest RMSE and Joint retains the lowest relative L2 and fRMSE.
Figure 3: Foil: sensitivity to simulation–experiment evaluation weighting. Rows show MORPH-Ti and MORPH-S and columns show RMSE, relative L2 , and fRMSE (lower is better). Joint gives the lowest error for all three metrics over λ=0.1 – 0.9 for MORPH-Ti and λ=0.1 – 0.8 for MORPH-S. Sim → Exp is best at the purely experimental endpoint ( λ=0 ), while Simulation only becomes best near the purely simulation endpoint, at λ=1 for MORPH-Ti and λ≥0.9 for MORPH-S.
Figure 4: FSI: sensitivity to simulation–experiment evaluation weighting. Rows show MORPH-Ti and MORPH-S and columns show RMSE, relative L2 , and fRMSE (lower is better). Joint gives the lowest error for all three metrics and both capacities throughout λ=0.2 – 0.9 . At λ=0.1 , Joint remains best for RMSE and fRMSE while Sim → Exp has the lowest relative L2 . Sim → Exp is best at λ=0 ; at λ=1 , Simulation only is best for MORPH-Ti, whereas Joint remains best for MORPH-S.
Sim → Exp
Joint
RMSE reduction (%) ↑
Dataset
Model
RMSE ↓
Rel. L2↓
fRMSE ↓
RMSE ↓
Rel. L2↓
fRMSE ↓
Cylinder
MORPH-Ti
0.292374
0.973175
0.050192
0.055293
0.179063
0.008347
81.1
MORPH-S
0.261182
0.869020
0.040659
0.034885
0.112863
0.005319
86.6
Controlled Cylinder
MORPH-Ti
0.262821
0.716707
0.067997
0.116303
0.315430
0.023091
55.7
MORPH-S
0.225042
0.613608
0.050295
0.071026
0.192083
0.013899
68.4
FSI
MORPH-Ti
0.661982
0.947801
0.102268
0.270900
0.400808
0.042135
59.1
Appendix
Table 9: Numerical-pressure retention after experimental adaptation. Pressure p is supervised in simulation but is unavailable in the experimental measurements. The table therefore focuses on Sim → Exp and Joint. Experiment only is not reported because pressure is never supervised in that protocol. Lower is better for the error metrics. The final column reports the percentage reduction in pressure RMSE achieved by Joint relative to Sim → Exp, 100(ESim→Exp−EJoint)/ESim→Exp . Joint is lower than Sim → Exp in all 24 metric comparisons (8 settings × 3 metrics).
Suppose a planner has a pre-trained simulator of a sequential decision problem and the option to run real experiments in the field. The simulator is cheap to query but inherits confounding and drift from its calibration data. Experimentation is unbiased but consumes one real unit per trial. We study when, and how, the planner should supplement the simulator with experiments. We give three results. First, an extended simulation lemma decomposes the simulator's value error into a calibration--deployment shift that randomization can identify and a parametric residual that no further interaction can reduce. Second, the value gap between the simulator-optimal policy and the optimum splits into a local component, on states the deployed policy already visits, and a reachability component, on states it does not. The reachability component stays bounded away from zero at any horizon under purely passive learning. Third, we propose Fisher-SEP, a simulation-aided experimental policy (SEP) that minimizes the posterior predictive variance of a target policy's value, with reward-only and transition-only specializations. Two case studies illustrate the regimes. In a vending-machine supply chain, front-loaded experimentation overtakes posterior updating once the horizon is long enough to amortize the pilot. In an HIV mobile-testing example with a corridor that separates a well-surveilled region from a poorly-surveilled one, only designed exploration reaches the poorly-surveilled region.
Harsh Parikh, Gabriel Levin-Konigsberg, Dominique Perrault-Joncas +1
Amazon SCOT, Seattle, USA · Yale University, New Haven, USA · Duke University, Durham, USA
Simulation-to-reality transfer, often called sim-to-real transfer, is a central challenge in robot learning. Yet, the tradeoff between measuring a system more accurately and training over a broader range of simulated dynamics is still poorly understood. In this work, we focused on the allocation of real-robot measurement time between system identification and domain randomization. We studied this tradeoff in a controlled sim-to-sim pendulum setting, where a hidden-parameter model stands in for the physical robot, and the experiment sweeps identification rollouts against the width of the randomization distribution. Across the reality gaps and noise levels we tested, the measurement budget did most of the work. A small number of identification rollouts closed most of the transfer gap, and once any real data was available, policies performed best when trained at the estimated parameters rather than over a widened randomization band. Broad randomization that contained the true system still did not substitute for measurement. These results hold in a benign regime where the dynamics are identifiable and only two parameters are unknown, so structural model mismatch remains the setting where randomization breadth may become more valuable. Overall, our results suggest that sim-to-real pipelines should first measure the parameters they can and reserve randomization for the uncertainty that remains.
Simulation can expand scarce real demonstrations for co-training, yet how world fidelity and similarity to human behavior affect policy performance remains unclear. We distinguish world grounding, which aligns simulation with the real system, and behavior grounding, which aligns simulated trajectories with human motion. We build a real2sim2real pipeline that varies these axes independently to generate data for co-training. On a dynamic dexterous pick-and-sort task, fully grounded co-training raises success from 52% to 86%; averaged across configurations, world grounding improves success by 18 percentage points and behavior grounding by 10. Deployed policies behave like a mixture of real-derived and simulation-derived policies, imitating real demonstrations in covered states and relying on simulated behavior elsewhere, which we examine through latent-space analysis. Together, these results suggest complementary roles: world grounding lets policies use simulated experience beyond real-data coverage, while behavior grounding matters mainly when world grounding is imperfect. Grounded simulation remains beneficial when co-training foundation models.