Intervention anchors and scientific verification in synthetic vascular predictive representations
Authors: Lingsen You, Yujun Guo, Xinyu Zhong, Zisu Peng, Wentong Wang, Li Shen, Junbo Ge
Organizations: Department of Cardiology, Zhongshan Hospital, Fudan University, Shanghai Institute of Cardiovascular Diseases, State Key Laboratory of Cardiovascular Diseases, NHC Key Laboratory of Ischemic Heart Diseases, National Clinical Research Center for Interventional Medicine, Key Laboratory of Viral Heart Diseases, Chinese Academy of Medical Sciences, Shanghai 200032, China. · Institutes of Biomedical Sciences, Fudan University, Shanghai 200030, China. · College of Biomedical Engineering, Fudan University, Shanghai 200438, China. · Department of Cardiology, Shidong Hospital, Yangpu District, Shanghai 200438, China.
Complete orthogonal predictive coordinates do not by themselves bind a latent direction to a named intervention. We present a mathematical and synthetic audit motivated by vascular device-vessel suitcordance. Capacity-matched least-squares predictors were exactly equivalent under complete fixed output transforms, whereas an anchor-only observer recovered interpretations only within the span of known perturbation signatures. Six three-dimensional configurations across 64 seeds gave a maximum paired prediction discrepancy of 6.7e-15 but a median untransported edit error of 1.513. Coordinate transport removed that error. Noisy and weak anchors constrained calibration stability, and changing the representation basis required recalibration or verified transport. Across 256 additional fits in dimensions 3-24, prediction equivalence persisted within 4.0e-15. We then evaluated nine deliberate runnable fault classes across 64 seeds. All 576 faulty executions completed, but each violated at least one reconstruction, prediction, delivered-edit or scope contract; all 320 valid control records passed. Repeating a faulty implementation gave exact self-agreement despite error against the separately computed simulator expectation. For one omitted-direction defect, probe coverage followed its analytic law, and rank-aware abstention protected unsupported interpretations. Scalar-noise experiments exposed both missed weak faults and excessive rejection under narrow relative tolerances. These controls provide an executable separation of prediction, semantic support and scientific acceptance. They are synthetic numerical audits, not clinical validation, neural JEPA-Anything replication, agent learning or patient treatment-effect estimation.
Figures & tables
Repeats
Inverse error
Procrustes error
1
0.324 (0.197–0.725)
0.172 (0.060–0.354)
4
0.155 (0.090–0.279)
0.086 (0.031–0.192)
16
0.084 (0.042–0.120)
0.045 (0.013–0.075)
64
0.039 (0.025–0.064)
0.019 (0.007–0.043)
Table 1: Full-rank anchor calibration at measurement-noise SD 0.2
Figure 1: Audit workflow . The upper path represents complete orthogonal analysis and synthesis. The lower path adds known perturbation signatures to calibrate interpretation. The three dimensionless state coordinates are conceptual analogues, not measured biological quantities. The audit evaluates coverage, noise sensitivity and coordinate transfer.
Figure 2: Prediction and semantic edits under basis rotation . A, the squared norm of the fixed residual (0.2, − 0.1, 0.3) remains 0.14 in either space. B, the error of an uncalibrated unit edit in the first coordinate increases with a rotation in the first two coordinates; transported edits recover the intended result. C, all factor SDs remain one for the separately constructed isotropic tight-frame control. This activity result is restricted to that control and does not hold for general covariance.
Figure 3: Anchor coverage and measurement noise . A, k independent coordinate anchors support k of three specified unit edits; unsupported components require abstention. This counting example uses orthogonal coordinate anchors. B–C, full-rank calibration errors for the unrestricted inverse and orthogonal Procrustes map, respectively. Curves show medians across 64 seeds; shaded bands show empirical 2.5th–97.5th percentiles. Increasing repeat counts are nested within each seed. Bands are descriptive simulation variation, not clinical confidence intervals.
Figure 4: Conditioning and coordinate transfer . A, full-rank anchors A = diag(1,s,s) become weaker as the true condition number increases. Measurement-noise SD is 0.02, averaged over 16 repeats. Median and empirical percentile bands summarize 64 seeds. B, each point denotes one independently sampled source–target basis pair. A source semantic map fails under target coordinate changes; noiseless target recalibration restores consistency. This test does not represent physiological domain shift.
Figure 5: Complete coordinates across synthetic dimensions . A, paired prediction discrepancies for 64 seeds in each of d = 3, 6, 12 and 24, using 256 training and 128 independent test observations in every dimension. Both predictors have d(2d + 1) parameters and d scalar output heads. B, untransported edit error; points are medians and bars empirical 2.5th–97.5th percentiles. C, unsupported fraction of d unit-coordinate edits with ⌊ d/2 ⌋ coordinate anchors. The supported-span identity was checked separately in each replicate. These are implementation stress tests, not measured vascular mechanisms.
Figure 6: Executable contracts and deterministic correction . A, passing fractions across 64 seeds for 14 runnable variants. Nine faulty variants give 576 faulty records; the healthy reference and four valid alternatives give 320 valid records. Five acceptance contracts assess execution, named-space round trip, noiseless prediction, actual delivered anchor edits and declared scope. The combined check requires all five. Consistent unit changes preserve physical outputs despite scaled latent coordinates. Another 64 execution-crash controls are recorded separately and omitted from the matrix. B, seed-zero untransported-edit control: prediction passes, the actual edit fails, prescribed coordinate transport restores the edit, and all checks pass on re-execution. The dashed line is the 10 -10 noiseless tolerance; plotting uses a floor of 10 -16 . This is a deterministic correction, not a learned agent repair. Fault counts describe a constructed bank, not natural error prevalence.
Figure 7: Independent probes of an unsupported direction . A, a sign reversal in the third direction leaves the measured rank-two anchor residuals near zero; the third direction is not anchored and has no fit residual. Independent axis probes reveal an error of two there. B, 4,096 nested random probe sequences sample axes uniformly with replacement. Detection follows 1 − (2/3) n for this single faulty axis. A single complementary-axis probe is a deterministic control with additional information. The original rank-two scope gate can withhold third-direction interpretation without detecting the defect. These probe budgets describe known noiseless simulator responses, not clinical detection rates.
Figure 8: Acceptance tolerance under scalar measurement noise . A, rejection of the constructed sign-reversal fault. B, erroneous rejection of healthy outputs. Healthy measurements are δ + ε and faulty measurements −δ + ε , compared with reference δ , with scalar ε∼ N(0,10 -8 ). At each of 25 δ values from 10 -6 to 1, 16,384 measurements per condition evaluate τ = 5 σ , τ = 0.05 δ and τ = max(5 σ ,0.05 δ ). Lines show analytic Gaussian probabilities and points the Monte Carlo rates. C, the three fixed thresholds relative to σ = 10 -4 . Coincident rules yield overlapping curves. The known response scale is not a clinical effect threshold; weak faults can be missed and narrow relative tolerances can reject healthy measurements.
Can a scientific agent distinguish a law it inferred from evidence from one it merely recognizes? We introduce Synthetic Universes, a controlled benchmark that pairs canonical famous worlds with matched twisted twins governed by nearby noncanonical mechanisms. We evaluate each reported law twice: by executing it on held-out continuations and transfer settings, and by independently checking whether it recovers the generating mechanism. In the current checkpoint of a pre-specified 60-cell study, 22 trials were graded and one additional run ended in infrastructure failure. Among 20 twin trials, 8 pass predictive verification while 5 recover the generator. The dissociation is bidirectional: six parsable outputs predict successfully while missing the mechanism, whereas three recover the mechanism but fail predictive rollout. Drag exhibits the first pattern (5/5 predictive pass, 1/5 mechanism recovery); Gravity exhibits the second (1/5 predictive pass, 4/5 mechanism recovery). Because matched famous controls, the corrected identifiability sweep, and the Evidence Ladder remain incomplete, we do not claim a confirmatory causal prior-conflict effect. Instead, the completed runs establish a narrower verification result: predictive adequacy and mechanism recovery are distinct scientific claims and require distinct tests.
Can learned state-action proposals exist in the physical world? Before executing a proposed action sequence, a robot can inspect empirical variation and disagreement with a learned transition model. However, aggregating these diagnostic signals obscures whether a proposal is inconsistent with the predictor or merely departs from recorded behavior. We formalize this prediction-control interface, separate the scalar trigger from its channel-wise diagnostic log, and establish that an all-pairs displacement term is redundant within a maximum that already contains the corresponding one-step term. We evaluate the monitors on 700 nominal and 5,250 synthetically perturbed 32-transition PushT windows, observing only planar pusher positions and goals. The transition-RMSE baseline achieves a ROC-AUC of 0.982, compared with 0.957 for a heterogeneous maximum and 0.972 for a spread-scaled residual.
Scientific machine learning is limited less by model size than by the data it is trained on. Observational data records what happened but not why; template synthetic data has a known generating process but only for the simulator's template, not the case a user faces. We argue a third option is now operationally feasible: instrumented data, in which every datum carries the mechanistic model that produced it, an explicit uncertainty over that model, and an executable family of counterfactuals. Verification-and-validation (V&V) instrumented image-to-simulation pipelines are one realisation: a sensor observation becomes a fully specified, solver-backed simulation with explicit, editable parameters and a propagated aleatoric/epistemic uncertainty. The substrate is case-specific, mechanistically supervised, and supports causal interventions through Pearl's do-operator. Near-term consequences for validation, auditing, and surrogate training span computational biology, climate, materials, fluid mechanics, and medical imaging; a longer-term, falsifiable implication concerns foundation models for scientific reasoning.
Daniel N. Wilke
School of Mechanical, Industrial and Aeronautical Engineering, University of the Witwatersrand, Johannesburg, South Africa