Can AI Scientists Change Their Minds? Prior-Evidence Conflict in Synthetic Universes
Organizations: University of California, Santa Cruz
Abstract
Can a scientific agent distinguish a law it inferred from evidence from one it merely recognizes? We introduce Synthetic Universes, a controlled benchmark that pairs canonical famous worlds with matched twisted twins governed by nearby noncanonical mechanisms. We evaluate each reported law twice: by executing it on held-out continuations and transfer settings, and by independently checking whether it recovers the generating mechanism. In the current checkpoint of a pre-specified 60-cell study, 22 trials were graded and one additional run ended in infrastructure failure. Among 20 twin trials, 8 pass predictive verification while 5 recover the generator. The dissociation is bidirectional: six parsable outputs predict successfully while missing the mechanism, whereas three recover the mechanism but fail predictive rollout. Drag exhibits the first pattern (5/5 predictive pass, 1/5 mechanism recovery); Gravity exhibits the second (1/5 predictive pass, 4/5 mechanism recovery). Because matched famous controls, the corrected identifiability sweep, and the Evidence Ladder remain incomplete, we do not claim a confirmatory causal prior-conflict effect. Instead, the completed runs establish a narrower verification result: predictive adequacy and mechanism recovery are distinct scientific claims and require distinct tests.
Figures & tables
| Family | Famous mechanism | Twisted mechanism |
|---|---|---|
| Spring | , | , |
| Gravity | , | , |
| Drag | , | , |
| Pendulum | ||
| Conservation † | audited target invariant class | generator under audit |
| Coupling † | audited pair-interaction class | generator under audit |
| Predictive pass | Mechanism recovery | |||
|---|---|---|---|---|
| Family | Famous | Twin | Famous | Twin |
| Spring | 1/1 | 0/5 | 1/1 | 0/5 |
| Gravity | 1/1 | 1/5 | 1/1 | 4/5 |
| Drag | – | 5/5 | – | 1/5 |
| Pendulum | – | 1/2 | – | 0/2 |
| Conservation | – | 1/2 | – | 0/2 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Artifact | Audit question |
|---|---|
| Task + observations | What evidence was actually available to the agent? |
| Raw response + tool trace | Which hypotheses and calculations were externally visible? |
| Predictive grader output | Did the reported law pass continuation and transfer? |
| Mechanism-checker output | Did the executable law match the generating equivalence class? |
| Parser regression tests | Could notation handling change a scientific label? |
| Failure logs | Was a missing result scientific, representational, or infrastructural? |
| Family | Scientific distinction | Audit quantity |
|---|---|---|
| Spring | restoring response vs. | matched power ; functional residual |
| Gravity | central-force exponent vs. | matched exponent ; functional residual |
| Drag | velocity exponent vs. | matched exponent ; functional residual |
| Pendulum | sinusoidal response vs. amplitude-dependent correction | fixed-domain functional residual / shape fit |
| Conservation | invariant equivalence class | normalized structural/equivalence residual |
| Coupling | pairwise interaction structure | product-vs.-alternative structural check |
| Claim | Current support | What would weaken or falsify it |
|---|---|---|
| Predictive adequacy and mechanism recovery are distinct verification targets | Both off-diagonal quadrants are populated; Drag and Gravity show replicated opposite dissociations. | A corrected evaluator that makes the off-diagonal cases disappear, or evidence that their labels arise from parser/grader artifacts. |
| Paired famous/twin worlds provide a controlled way to study prior–evidence conflict | The observation interface is shared while the hidden mechanism changes. | A demonstrable cue that reveals world identity, or contamination/leakage that gives access to generator labels. |
| The present data establish the magnitude of a prior-conflict penalty | Not claimed. Famous controls are incomplete. | Requires completion of the matched grid and the pre-specified paired analysis. |
| Increasing discriminative evidence causes prior abandonment | Not claimed. The Evidence Ladder is incomplete. | A flat or reversed recovery trend under increasing complexity-adjusted oracle evidence would directly challenge this interpretation. |
| The observed twin failures are specific to pretrained language-model priors | Not claimed. No prior-free or cross-model baseline is complete. | A symbolic baseline showing comparable failures, or strong dependence on agent architecture, would weaken a pretraining-prior explanation. |
| The uniform rollout threshold fully characterizes scientific correctness | Not claimed. Rollout and mechanism are deliberately separate. | Large agreement on derivative/geometric diagnostics despite raw rollout failure would show that pointwise long-horizon NRMSE is too coarse as a standalone verifier. |