Reliable world models should not only predict future states but express how actions change the world in an explicit, transparent and testable form, such as equations. Yet methods that rely on a fixed set of trajectories cannot distinguish equally good competing hypotheses, while searches over a fixed set of predefined candidates cannot discover equations outside the initial hypothesis space. We introduce ALDER (Action-guided Law Discovery, Evaluation, and Revision), a method that actively proposes novel experiments to test and revise models. Specifically, ALDER proposes parametric equations; a numerical optimizer fits their coefficients; an independent verifier tests these candidates on held-out data. To distinguish between competing valid hypotheses, a cost- and safety-aware selector queries interventions, in the form of novel experiments. The resulting counterexamples update the evidence ledger and guide the next structural revision, while incompatible laws are discarded. Across an in-house benchmark, ODE equation discovery tasks, and robotic experiments, ALDER discovers laws beyond its initial formula set, repairs failed model proposals, distinguishes fixed candidate models with fewer interactions, and improves out-of-distribution prediction. Furthermore, given a current state and a target, ALDER selects control actions by solving the inverse problem defined by its validated world model. Together, these results show that explicit equation-based world models can be tested and revised through interaction, then naturally used to guide goal-directed control.
Figures & tables
Figure 1: Different world models can fit the same trajectories; Alder selects experiments to distinguish them. At mass m0 , both equations fit the observations (A). At 2m0 with the same force, their predicted accelerations differ. The observation contradicts the mass-independent prediction (B).
Figure 2: Alder uses new observations to revise equations. The Law Revision Agent proposes an equation structure from public evidence, and the Coefficient Fitter estimates its continuous parameters. The Protected Admission Verifier checks fitted equations on held-out data; admitted equations can be used for prediction and action selection. After rejection, ARS may select a feasible experiment where candidate equations predict different outcomes; its observation guides the next revision.
A. Controlled discovery
NovelLaw (60 cases)
PullCubeTool (5 runs)
PushT (5 runs)
Method
Functional recovery (%, ↑ )
Shift NMSE ( ↓ )
Admitted ( ↑ )
Shift NMSE ( ↓ )
Admitted ( ↑ )
Shift NMSE ( ↓ )
One-shot agent
43.3
6.54×10−2
0/5
4.816±1.973
3/5
0.842±0.644
Revision-only agent
75.0
1.50×10−2
0/5
1.296±0.024
3/5
1.292±1.638
Random-query agent
85.0
9.15×10−3
5/5
0.194±0.138
4/5
0.673±0.473
Alder
85.0
5.58×10−3
5/5
0.161±0.052
4/5
0.420±0.185
Table 1: Interactive revision improves functional recovery; fitted equations predict ODEBase trajectories under shift. A: NovelLaw reports functional recovery (%) and median shifted NMSE over 60 cases; robot shifted NMSE is mean ± SD over five episode-disjoint partitions. Rejected runs are included. B: median trajectory NMSE and mean symbolic accuracy on 59 ODEBase systems.
OpenDrawer
OpenDoor
Model
OOD prediction RMSE ( ↓ )
Control MAE ( ↓ )
Success (%, ↑ )
OOD prediction RMSE ( ↓ )
Control MAE ( ↓ )
Success (%, ↑ )
Nominal
6.8±0.7
5.04±0.54
83.6±2.2
7.2±0.8
4.26±0.37
94.4±2.0
Passive law
7.5±1.0
5.78±0.76
80.3±3.6
7.4±1.5
3.80±0.49
89.7±3.1
Matched MLP
3.7±0.8
1.98±0.35
96.9±2.4
3.1±0.7
1.81±0.32
99.5±1.1
ALDER
0.9±0.1
0.74±0.10
100.0±0.0
0.9±0.1
0.49±0.06
100.0±0.0
Table 2: Explicit dynamics discovered by Alder improve shifted prediction while preserving downstream control performance in RoboCasa. Prediction RMSE and control MAE are in 10−3 normalized joint-position units; success rates are percentages. Values are mean ± SD across ten seeds, each averaging four mechanism families. Control uses 640 targets per model and task; bold denotes the best reported performance.
A. Missing structure or state
Regime
Admitted models
Localization AUROC ( ↑ )
Gated NMSE ( ×10−3 , ↓ )
Complete
100.0%
–
0.042±0.029
Missing smooth
0.0%
0.494±0.019
14.5±3.0
Missing local
0.0%
0.969±0.007
18.5±2.9
Hidden local
0.0%
0.997±0.001
168±114
Table 3: Admission checks reject models with missing equation terms or state variables; degraded force observations reduce prediction coverage. A: AUROC and gated-residual NMSE are evaluated under distribution shift and do not determine admission. B: Drawer and Door are pooled; coverage is the fraction of worlds with admitted equations. RMSE uses those worlds only, whereas control success covers all targets, counting refusals as failures. Values with ± are mean ± SD.
protected division, power, exponential, trigonometric functions, absolute value, maximum
angular displacement; 95 nodes
Door
current and initial hinge/latch angles
literal power, sine, cosine
displacement; 127 nodes
Peg
hole-frame x,y,z and hole radius
literal power, absolute value, pairwise maximum
signed score; 63 nodes
Appendix
Table 4: Typed grammars expose building blocks, not complete target equations. In RQ1, the variables and primitive operators are available, but complete target expressions and harness-held reference bases are absent from the proposal sandbox.
NovelLaw (60 cases)
PullCubeTool (5 runs)
PushT (5 runs)
Method
Functional recovery (%, ↑ )
Shift NMSE ( ↓ )
Admitted ( ↑ )
Shift NMSE ( ↓ )
Admitted ( ↑ )
Shift NMSE ( ↓ )
Baselines
Fixed/simple model
0.0
5.94×10−1
0/5
5.773±0.287
0/5
2.762±1.098
Polynomial SINDy
28.3
1.16×10−1
0/5
16.573±10.539
0/5
18.281±18.708
Matched MLP
0.0
4.20×10−1
0/5
5.570±0.407
1/5
1.815±1.609
Variants
Appendix
Table 5: Closed-loop revision improves functional recovery and shifted prediction over passive baselines. The full controlled comparison supplements Table 1 A. NovelLaw reports median NMSE over 60 cases; robot errors are mean ± SD over five episode-disjoint partitions. Rejected runs are included. Bold marks the best result in each column, including ties.
Split
Structures
Cases
Structural role
Development
8
16
Covers every primitive component; used only for protocol and prompt development.
IID
4
12
Reuses a development motif with unseen coefficients and observations.
Compositional
12
36
Contains exactly one reserved component pairing.
Stress
4
12
Contains two simultaneous reserved component pairings.
Appendix
Table 6: NovelLaw tests new compositions of familiar primitives. Each test structure has three independently sampled parameter/data instances.
Method
IID
Comp.
Stress
All
Shift NMSE
Queries
One-shot
41.7
50.0
25.0
43.3
0.0654
0
Revision-only
58.3
86.1
58.3
75.0
0.0150
0
Random-query
91.7
91.7
58.3
85.0
0.00915
24
Alder
91.7
91.7
58.3
85.0
0.00558
24
Appendix
Table 7: Revised proposals recover more equations than one-shot proposals. RQ1 values are functional recovery (%); queries are medians over all 60 cases. Independent model calls make the last three rows an auxiliary sensitivity analysis, not a controlled selector ablation.
Comparator
Median shift ratio
Wilcoxon p
Recovery wins/losses
One-shot agent
0.0782
5.18×10−9
25/0
Revision-only
0.146
2.66×10−7
7/1
Random-query
0.901
0.203
3/3
Fixed library
0.00782
1.63×10−11
51/0
Polynomial SINDy
0.0249
2.32×10−10
34/0
MLP
0.0130
1.80×10−11
51/0
Appendix
Table 8: Paired tests support gains over one-shot and revision-only discovery. RQ1 shift-error ratios below one favor closed-loop Alder .
Model and evidence
PullCubeTool
PushT
Linear, public
5.773
2.762
Polynomial SINDy, public
16.573
18.281
MLP, public
5.570
1.815
Supplied structure, public
4.579
0.919
Supplied structure, all queries
0.212
0.588
Appendix
Table 9: Query evidence improves prediction even when structure is supplied. Hard-track entries are mean shifted NMSE over five folds; lower is better. “All queries” uses every initially unrevealed query outcome and is not an Alder discovery result.
PullCubeTool
PushT
Method
Admitted
Shift NMSE
Shift MAE
Admitted
Shift NMSE
Shift MAE
One-shot
0/5
4.816±1.973
1.081±0.368
3/5
0.842±0.644
0.086±0.032
Revision-only
0/5
1.296±0.024
0.377±0.032
3/5
1.292±1.638
0.106±0.072
Random-query
5/5
0.194±0.138
0.144±0.064
4/5
0.673±0.473
0.078±0.021
Alder
5/5
0.161±0.052
0.136±0.029
4/5
0.420±0.185
0.071±0.024
Appendix
Table 10: Interactive revision improves model admission on both robot tracks. Values are mean ± standard deviation over five folds. MAE is in centimeters for PullCubeTool and radians for PushT.
Acquisition
Identified ( ↑ )
Mean cost † ( ↓ )
Median cost † ( ↓ )
Passive sampling
4.4%
8.85
9
Random selection
94.0%
3.53
3
Geometric coverage
100.0%
1.91
1
Parameter uncertainty
99.5%
1.35
1
Alder (ARS)
100.0%
1.16
1
Oracle reference
100.0%
1.18
1
Appendix
Table 11: ARS identifies all targets with the lowest mean penalized cost. RQ2 uses 182 initially ambiguous cases, the same candidate models, and an eight-query budget. Bold marks the best non-oracle results, including ties.
Comparator
Mean cost reduction [95% CI]
Win/tie/loss
Passive
7.69 [7.55, 7.80]
182/0/0
Random
2.36 [2.05, 2.68]
151/29/2
Coverage
0.74 [0.56, 0.94]
63/118/1
Parameter uncertainty
0.18 [0.08, 0.32]
18/161/3
Appendix
Table 12: ARS reduces mean identification cost relative to each listed comparator. Paired RQ2 reductions are positive when they favor ARS.
Task
Selector
Identified
Support
Median queries
Safety
Drawer
Passive
50.0%
50.0%
3.5
0
Random
100.0%
97.5%
0.5
0
Coverage
100.0%
70.0%
0.5
0
Uncertainty
100.0%
100.0%
0.5
0
ALDER
100.0%
97.5%
0.5
0
Door
Passive
25.0%
25.0%
7.0
0
Appendix
Table 13: Structured queries identify Door models sooner than random queries. RoboCasa candidates are fixed. Queries are censored at seven after exhausting the six-query budget. “Support” reports exact recovery of the optional equation terms after coefficient thresholding.
OpenDrawer
OpenDoor
Model
OOD prediction RMSE ( ↓ )
Relative success (%, ↑ )
OOD prediction RMSE ( ↓ )
Relative success (%, ↑ )
Nominal
0.0068±0.0007
96.7±1.4
0.0072±0.0008
95.9±2.0
Passive law
0.0075±0.0010
94.2±2.1
0.0074±0.0015
96.4±2.4
Matched MLP
0.0037±0.0008
99.5±1.1
0.0031±0.0007
100.0±0.0
ALDER law
0.0009±0.0001
100.0±0.0
0.0009±0.0001
100.0±0.0
Parameter oracle
0.0009±0.0001
100.0±0.0
0.0009±0.0001
100.0±0.0
Appendix
Table 14: Equation-based models improve prediction under the original finite-action evaluation. Values are mean ± SD over ten seeds, each averaging four dynamics families. Relative success means executed target error within 0.01 of the best finite candidate’s error; it is not the absolute target-reaching success in Table 2 . Each task has 640 targets per model. Bold marks the best non-oracle results, including ties.
Regime
Equation only
MLP
Always residual
Gated residual
Complete
4.20×10−5
0.0109
0.000345
4.20×10−5
Missing smooth
0.217
0.0133
0.00380
0.0145
Missing local
0.0485
0.0308
0.0191
0.0185
Hidden local
0.108
0.134
0.183
0.168
Appendix
Table 15: Residual correction helps missing-term regimes but not hidden-state failures. RQ4 reports shifted prediction NMSE; lowest error is bold.
Method
Recon. NMSE ↓
Gen. NMSE ↓
OOD NMSE ↓
Complexity
Sym. Acc (%) ↑
Passive Symbolic Discovery
SINDy
5.07e-04
5.09e-01
1.41e+00
11.7
18.5
Operon
7.95e-05
2.57e-01
1.84e+00
14.0
2.4
PySR
2.84e-03
8.82e-01
1.43e+00
6.6
18.1
E2E
3.69e-01
1.49e+00
2.19e+00
52.6
0.0
ODEFormer
4.61e-03
3.83e-01
2.04e+00
15.9
16.5
Appendix
Table 16: Alder achieves low median error across ODEBench prediction settings. We report the median NMSE (lower is better) across reconstruction, generalization, and future-time OOD prediction, mean symbolic accuracy (higher is better), and mean expression complexity (ODEBench mean = 19.3). Best NMSE and symbolic-accuracy results are in bold, and second-best results are underlined. The Alder row is shaded.
Method
Public dev.
Protected ID
High-angle OOD
Admitted
Zero displacement
6.671
6.501
28.866
0%
Linear joints
0.810
0.813
11.660
0%
Polynomial SINDy-3
0.025
0.027
1.455
40%
MLP
0.047
0.051
10.546
0%
Additive trigonometric library
0.068
0.065
2.626
0%
Full composed reference
0.020
0.021
0.568
100%
Appendix
Table 18: One-shot joint equations extrapolate to unseen Door angles. Adroit Door uses five episode partitions. Errors are handle- x MAE in cm; lower is better. OOD contains only unseen high-angle states.
Method
Public dev.
Protected ID
Radius OOD
OOD success F1
Admitted
Always failure
0.500
0.500
0.500
0.000
0%
Linear
0.557
0.553
0.552
0.266
0%
Polynomial logistic-3
0.769
0.767
0.754
0.450
0%
MLP
0.905
0.901
0.796
0.605
0%
Additive absolute rule
0.649
0.646
0.638
0.369
0%
Full non-smooth reference
1.000
1.000
1.000
1.000
100%
Appendix
Table 19: One-shot constraints generalize across held-out peg-hole radii. PegInsertionSide provides binary simulator labels. All metrics except “Admitted” are balanced accuracy or success F1; higher is better.
Backbone
Total/active B
Admitted
Support
Queries
Pred. cover
NMSE
Tokens
GPT-5.6 (primary)
not public
100.0%
83.3%
0.5
100.0%
7.85×10−8
345788
Qwen3.5-9B
9/9
100.0%
66.7%
1.0
100.0%
3.34×10−7
35082
GPT-OSS-20B
21/3.6
100.0%
41.7%
1.0
100.0%
3.16×10−6
43037
Qwen3.8-27B
27/27
91.7%
75.0%
0.5
100.0%
8.77×10−7
56292
Qwen3.6-35B-A3B
35/3
91.7%
41.7%
1.0
100.0%
3.93×10−6
54832
Qwen3-Coder-Next
80/3
58.3%
33.3%
1.5
100.0%
5.85×10−5
78094
Appendix
Table 20: Alder discovers validated equations with multiple open-weight backbones. On 12 matched RoboCasa worlds, the first row is the primary backend; the remaining seven are open-weight ablations. Queries are medians with non-admitted runs assigned budget plus one; NMSE is the geometric mean over finite predictions. Total/active parameter counts follow public model declarations, while the primary parameter count is not public. Tokens are API-reported mean request totals and are not a quality-normalized cost metric. “Support” uses the same strict optional-term recovery criterion as the primary study.
Method
Calls
Proposals
Queries
Input tokens
Agent h
Revision-only
60
138
0
31.8M
1.97
Random-query
60
117
1,368
30.9M
2.15
Alder
60
125
1,560
34.1M
2.22
Appendix
Table 21: Each full agent variant uses one invocation per case. RQ1 usage is summed over 60 cases. One-shot is derived from Alder ’s stored first round and incurs no additional call.