Spectra: Exact Component Transport for Test-Time Prior Adaptation in Simulation-Based Inference
Authors: Xin Zhao, Nico Scherf, Robert Trampel, Kerrin J. Pine, Nikolaus Weiskopf
Organizations: Department of Neurophysics, Max Planck Institute for Human Cognitive and Brain Sciences, Leipzig, Germany · International Max Planck Research School on Cognitive NeuroImaging, Leipzig, Germany · Methods and Development Group Neural Data Science and Statistical Computing, Max Planck Institute for Human Cognitive and Brain Sciences, Leipzig, Germany · Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI), Dresden/Leipzig, Germany · Felix Bloch Institute for Solid State Physics, Faculty of Physics and Earth System Sciences, Leipzig University, Leipzig, Germany · Functional Imaging Laboratory, Department of Imaging Neuroscience, UCL Queen Square Institute of Neurology, University College London, London, UK
Simulation-based inference (SBI) has become a powerful approach to Bayesian inference in complex scientific models whose likelihoods are difficult or impossible to evaluate. Amortized SBI learns reusable inference models from simulated data, enabling rapid posterior inference for new observations, and modern generative models have made these models increasingly expressive. However, this reuse is limited to the prior distribution chosen during training, whereas scientific analyses often need revised priors as knowledge accumulates or alternative assumptions are tested. We introduce Spectra, a test-time adaptation method for diffusion-based SBI. Spectra uses an exact score-transport identity to obtain the adapted score from a frozen diffusion model in closed form for structured prior changes, without additional simulation or training. Across six SBI benchmarks, Spectra achieves accurate adaptation under strong prior shifts at low online sampling cost. This enables pretrained SBI models to incorporate updated prior information at test time.
Figures & tables
Figure 1: Prior shift and test-time adaptation on the simple likelihood complex posterior (SLCP) task. (a) A concentrated target prior replaces the broad training prior. (b) The frozen model represents the base posterior under the training prior, while the target prior concentrates posterior mass on two of its four modes. (c) Both PriorGuide and Spectra recover the target modes. However, PriorGuide retains visible mass outside the target contours, whereas Spectra follows the target geometry more closely. The posterior panels show the (θ3,θ4) projection. Full results across tasks and observations are reported in Table 1 and Figure 3 .
Single-factor, mild
Single-factor, strong
Mixture
Task
Method
C2ST
MMTV
C2ST
MMTV
C2ST
MMTV
Two Moons
Base
0.568 (0.036)
0.219 (0.087)
0.691 (0.095)
0.426 (0.192)
0.736 (0.074)
0.510 (0.147)
SIR
0.525 (0.005)
0.075 (0.011)
0.544 (0.018)
0.070 (0.012)
0.543 (0.018)
0.062 (0.009)
PriorGuide
0.528 (0.012)
0.123 (0.046)
0.556 (0.034)
0.165 (0.086)
0.561 (0.036)
0.176 (0.086)
Spectra
0.521 (0.005)
0.093 (0.016)
0.532 (0.015)
0.102 (0.030)
0.521 (0.009)
0.073 (0.014)
SLCP
Base
0.765 (0.057)
0.255 (0.041)
0.917 (0.019)
0.466 (0.072)
0.946 (0.029)
0.498 (0.075)
Table 1: Posterior accuracy under single-factor and mixture prior shifts. Mean (standard deviation) across ten observations or 30 prior–observation pairs. Ideal values are 0.5 for C2ST and 0 for MMTV. Bold marks the mean closest to the ideal value within each task and setting.
Figure 2: Accuracy gap as the target prior narrows. Mean paired PriorGuide–Spectra difference with 95% paired t confidence intervals across ten observations. All experimental factors except target-prior width are held fixed. Positive gaps indicate better performance by Spectra.
Figure 3: Posterior geometry under strong single-factor prior shifts. Reference contours enclose 50% and 90% posterior mass. Markers show 1,000 samples from each method for one fixed observation, backbone, and sampling seed per task. Larger SIR markers indicate repeated resampling of the same Base sample. SLCP and GL-20D show two-dimensional projections. Table 1 reports full-dimensional accuracy.
Figure 4: Single-factor accuracy versus online sampling time under strong prior shifts. PriorGuide (VJP) uses a vector–Jacobian product for the same guidance update and preserves PriorGuide’s accuracy (Appendix M ). Timing covers the reverse sampling stage for 1,000 posterior samples after compilation. One-time setup costs are excluded.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
πtr,πnew,r
training prior, target prior, and their ratio
p,q
base and target posteriors
qϕ
posterior induced by a single prior-ratio factor
t , τ=σ(t)2
diffusion time and corresponding noise variance
pτ,sp,sp
Gaussian-smoothed base posterior, its exact score, and frozen score estimate
ϕk,bk,qk
prior-ratio factor, its coefficient, and posterior component
Appendix
Table 2: Principal notation. Symbols specific to error analysis and sampling protocols are defined where they are used.
Task
Reference
Validation
GL-10D/20D
closed-form posterior
analytic
Two Moons
exact-likelihood quadrature on a 4096×4096 grid
grid refinement changes the log normalizer by at most 2×10−4 nats
OUP
exact-likelihood quadrature on a 4096×4096 grid
grid refinement changes the log normalizer by less than 10−12 nats; target mass on the grid boundary at most 4×10−10 (both shifts)
SLCP
tempered SMC, 40,000 particles, four independent runs
cross-run C2ST at most 0.509 (strong) and 0.521 (mild)
BCI
tempered SMC, 8,000 particles, four independent runs; Gauss–Hermite likelihood of order 301
cross-run C2ST at most 0.509 (both shifts and mixtures); quadrature-order check
Appendix
Table 3: Posterior references and validation. Reported validation statistics are worst-case values over the evaluated single-factor observations; BCI also includes the mixture pairs. Cross-run C2ST compares independent reference runs of the same posterior (ideal value 0.5 ).
Mild
Strong
Task
Method
C2ST
MMTV
RMSE
C2ST
MMTV
RMSE
Two Moons
Base
0.568 (0.036)
0.219 (0.087)
0.397 (0.197)
0.691 (0.095)
0.426 (0.192)
0.174 (0.106)
SIR
0.525 (0.005)
0.075 (0.011)
0.340 (0.147)
0.544 (0.018)
0.070 (0.012)
0.100 (0.061)
PriorGuide (exact ratio)
0.522 (0.010)
0.105 (0.040)
0.371 (0.173)
0.554 (0.032)
0.167 (0.085)
0.136 (0.063)
PG-FullCov
0.528 (0.015)
0.123 (0.057)
0.379 (0.180)
0.552 (0.038)
0.159 (0.093)
0.142 (0.070)
Spectra
0.512 (0.006)
0.072 (0.007)
0.346 (0.151)
0.519 (0.009)
0.068 (0.012)
0.102 (0.059)
Appendix
Table 4: Single-factor comparison at (N,L)=(100,0) . Mean ( standard deviation ) across ten observations, after averaging sampling seeds within each backbone and then averaging the three backbones within each observation. C2ST is ideal at 0.5 and MMTV at zero. Bold indicates the displayed mean closest to the ideal value within each task and shift, including rounded ties. RMSE to the generating parameter is a supplementary diagnostic and is not bolded.
Shift
Task
PriorGuide − Spectra
PG-FullCov − Spectra
Mild
Two Moons
+0.0097[+0.0018,+0.0177]
+0.0151[+0.0039,+0.0264]
SLCP
+0.0019[−0.0070,+0.0108]
+0.0031[−0.0026,+0.0087]
OUP
+0.0014[−0.0025,+0.0052]
+0.0011[−0.0011,+0.0034]
BCI
+0.0009[−0.0022,+0.0040]
−0.0026[−0.0047,−0.0004]
GL-10D
+0.0113[+0.0044,+0.0182]
−0.0049[−0.0074,−0.0024]
GL-20D
+0.0309[+0.0130,+0.0488]
−0.0073[−0.0109,−0.0038]
Appendix
Table 5: Paired C2ST differences at (N,L)=(100,0) . Entries are baseline minus Spectra, with 95% paired Student- t confidence intervals across ten observations. Positive values correspond to lower C2ST for Spectra.
Two Moons
SLCP
GL-20D
Method
C2ST
MMTV
C2ST
MMTV
C2ST
MMTV
PriorGuide (exact ratio)
0.560 (0.039)
0.174 (0.099)
0.617 (0.029)
0.115 (0.030)
0.531 (0.025)
0.035 (0.004)
PG-FullCov
0.567 (0.047)
0.192 (0.118)
0.620 (0.030)
0.117 (0.032)
0.530 (0.023)
0.035 (0.004)
Spectra
0.534 (0.017)
0.103 (0.033)
0.586 (0.038)
0.100 (0.041)
0.531 (0.024)
0.035 (0.004)
Appendix
Table 6: Held-out single-factor comparison with Langevin corrections. Strong shifts at (N,L)=(25,8) . Mean ( standard deviation ) across seven held-out observations after averaging three sampling seeds within each of three backbones. C2ST is ideal at 0.5 and MMTV at zero. Bold indicates the displayed mean closest to the ideal value within each task, including rounded ties.
C2ST
MMTV
Task
Width
Difference [95% CI]
Favor
Difference [95% CI]
Favor
Two Moons
0.50
+ 0.0075 [ + 0.0006, + 0.0144]
10/10
+ 0.0297 [ + 0.0062, + 0.0532]
10/10
0.40
+ 0.0191 [ + 0.0032, + 0.0350]
9/10
+ 0.0596 [ + 0.0203, + 0.0990]
10/10
0.30
+ 0.0434 [ + 0.0179, + 0.0688]
10/10
+ 0.1140 [ + 0.0532, + 0.1748]
10/10
0.25
+ 0.0571 [ + 0.0280, + 0.0861]
9/10
+ 0.1486 [ + 0.0788, + 0.2185]
10/10
0.20
+ 0.0727 [ + 0.0414, + 0.1039]
10/10
+ 0.1828 [ + 0.1065, + 0.2590]
10/10
Appendix
Table 7: Fixed-center shift-severity sweep. Entries are mean paired differences (PriorGuide minus Spectra) across ten observations, with 95% paired Student- t confidence intervals after averaging over three sampling seeds and three backbones within each observation. Width is the target-prior marginal standard deviation relative to the training prior. Positive differences favor Spectra. Favor gives the number of observations with a positive paired difference.
Spectra
Task
Base
SIR
PriorGuide
PG-FullCov
Direct
PS
Ref
C2ST
Two Moons
0.736 (0.074)
0.543 (0.018)
0.552 (0.035)
0.538 (0.021)
0.514 (0.007)
0.515 (0.007)
0.515 (0.008)
SLCP
0.946 (0.029)
0.796 (0.086)
0.611 (0.043)
0.587 (0.042)
0.554 (0.040)
0.555 (0.040)
0.554 (0.040)
OUP
0.703 (0.063)
0.529 (0.011)
0.517 (0.009)
0.507 (0.008)
0.506 (0.006)
0.507 (0.006)
0.506 (0.006)
GL-10D
0.996 (0.003)
0.999 (0.001)
0.598 (0.035)
0.539 (0.025)
0.570 (0.094)
0.536 (0.023)
0.531 (0.025)
Appendix
Table 8: Mixture posterior inference across 180 prior–observation pairs. Mean ( standard deviation ) across 30 pairs per task, after averaging sampling seeds within each backbone and then the three backbones within each pair. All output samplers use (N,L)=(100,0) . PriorGuide receives the exact prior ratio. Direct, PS, and Ref are the Spectra variants of Appendix J.4 . Ref uses reference component weights and is included only as a diagnostic. C2ST is ideal at 0.5 and MMTV at zero. Bold indicates the displayed mean closest to the ideal value among methods that do not use reference weights, including rounded ties.
(100,0)
(25,8)
Method
C2ST
MMTV
C2ST
MMTV
PriorGuide
0.568 (0.043)
0.112 (0.038)
0.549 (0.038)
0.100 (0.044)
Spectra-Direct
0.591 (0.123)
0.158 (0.171)
0.595 (0.122)
0.161 (0.170)
Spectra-PS
0.529 (0.022)
0.072 (0.023)
0.534 (0.026)
0.074 (0.024)
Spectra-Ref
0.519 (0.019)
0.049 (0.012)
–
–
Appendix
Table 9: Mixtures with substantial mass in both posterior components. Mean ( standard deviation ) across the 18 pairs for which each component carries at least 5% reference posterior mass. Weight estimates are fixed across output configurations. Spectra-Ref uses reference component weights and is shown only as a diagnostic at (N,L)=(100,0) .
C2ST
MMTV
K
Spectra-PS
Spectra-Ref
PriorGuide
Spectra-PS
Spectra-Ref
PriorGuide
2
0.517 (0.008)
0.518 (0.009)
0.573 (0.062)
0.064 (0.016)
0.070 (0.011)
0.211 (0.152)
4
0.521 (0.013)
0.517 (0.008)
0.586 (0.049)
0.074 (0.024)
0.075 (0.011)
0.260 (0.134)
8
0.520 (0.010)
0.519 (0.009)
0.555 (0.049)
0.062 (0.004)
0.069 (0.004)
0.178 (0.141)
16
0.507 (0.001)
0.509 (0.003)
0.545 (0.038)
0.063 (0.011)
0.070 (0.002)
0.170 (0.115)
Appendix
Table 10: Accuracy as the number of prior components increases on Two Moons. Mean ( standard deviation ) across three observations at (N,L)=(25,8) after averaging three sampling seeds within each backbone. Spectra-PS uses one backbone, while Spectra-Ref and PriorGuide use three backbones. All methods are evaluated against the support-matched reference.
Reference weights
Path-space estimate
K
Effective components
Components with at least 5%
Effective components
Largest weight error
2
1.00 / 1.02 / 1.00
1 / 1 / 1
1.00 / 1.02 / 1.00
3.4×10−4
4
1.19 / 1.36 / 1.12
2 / 2 / 2
1.19 / 1.35 / 1.12
2.2×10−3
8
1.21 / 3.09 / 2.18
2 / 3 / 2
1.21 / 3.07 / 2.18
1.6×10−2
16
2.50 / 4.54 / 4.10
5 / 5 / 5
2.50 / 4.52 / 4.15
8.4×10−3
Appendix
Table 11: Component allocation and path-space weight accuracy. For each of the three observations, the table reports the effective number of components 1/∑kαk2 and the number carrying at least 5% of the posterior mass. The final column gives the largest absolute difference between path-space and reference component weights across all components and observations.
Figure 5: Online sampling time across the five timed tasks. Each task uses the sampler configuration of the main benchmark and the timing protocol above. Labels show the runtime ratio of PriorGuide (VJP) to Spectra.
Figure 6: Single-factor MMTV versus online sampling time. The figure adds PG-FullCov to the methods shown in Figure 4 . Filled circles mark (25,8) , squares mark (100,0) , and hollow circles mark the remaining sampler configurations. Error bars are observation-level standard errors. Both axes are logarithmic.
Figure 7: Cost of the first 1,000 mixture posterior samples. Stacks separate precomputation, component-weight estimation, and online sampling for PriorGuide, PG-FullCov, Spectra-Direct, and Spectra-PS. Labels give the total displayed cost. Panels use different linear time scales.
K
1
2
4
8
16
32
64
Spectra
0.1166
0.1165
0.1165
0.1165
0.1165
0.1165
0.1164
PriorGuide
0.4464
0.4440
0.4442
0.4446
0.4456
0.4466
0.4482
Appendix
Table 12: Online sampling time as the number of prior components increases on Two Moons. Median seconds for 1,000 posterior samples at (N,L)=(25,8) under the timing protocol of Appendix M .
γ
Orient.
K
SIR
PriorGuide
Spectra-Ref
Spectra-PS
C2ST
1
iso
1
0.613 (0.069)
0.621 (0.055)
0.548 (0.016)
0.548 (0.016)
2
+
52
0.589 (0.090)
0.621 (0.092)
0.528 (0.013)
0.570 (0.143) †
2
−
52
0.578 (0.044)
0.604 (0.063)
0.528 (0.012)
0.535 (0.027)
4
+
104
0.599 (0.070)
0.609 (0.084)
0.531 (0.023)
0.571 (0.149) †
4
−
104
0.590 (0.073)
0.600 (0.084)
0.527 (0.018)
0.530 (0.021)
Appendix
Table 13: Anisotropy sweep on Two Moons. Mean (standard deviation) over ten observations, scored against support-matched references for the exact target. Each observation’s value averages three backbones, and for PriorGuide and Spectra three sampling seeds per backbone. Sampling uses (N,L)=(25,8) , except (100,0) for SIR. At γ=1 , K=1 and Spectra-Ref and Spectra-PS coincide.
Observation
Approx. TV
PriorGuide
PG-FullCov
Spectra-Ref
Spectra-PS
1
0.00306
0.535±0.003
0.527±0.007
0.507±0.006
0.513±0.006
2
0.00456
0.519±0.011
0.519±0.013
0.518±0.008
0.518±0.006
3
0.00285
0.549±0.008
0.526±0.009
0.530±0.009
0.529±0.010
Appendix
Table 14: Correlated-prior approximation on Two Moons. Approx. TV is the posterior total variation distance induced by the K=64 finite quadrature, evaluated on a 4096×4096 grid. C2ST is mean ± standard deviation over three sampling seeds. Approx. TV quantifies posterior error from the finite prior approximation, while C2ST evaluates the sampled posterior against the exact target. PriorGuide and PG-FullCov use the exact ratio; Spectra-Ref and Spectra-PS use the quadrature approximation with reference and path-space component weights, respectively.
Simulation-based inference (SBI) provides amortized Bayesian parameter inference from simulator-generated data without requiring explicit likelihood evaluation. Its reliability can degrade under model misspecification, where real-world observations are not well represented by the simulator used for training. Existing methods using unlabeled real-world data often align simulated and real-world data distributions, but marginal alignment alone does not directly preserve parameter-relevant information needed for posterior inference. We propose SPIN, an SBI framework with parameter-relevant information-preserving domain transfer using unlabeled, unpaired real-world observations. During training, SPIN translates labeled simulator observations toward the real-world domain and back to the simulator domain, using the original simulator labels to encourage domain transfer that preserves parameter-relevant mutual information. At test time, the learned real-to-simulator transport maps real-world observations into the simulator domain for posterior inference, without requiring real-world parameter labels or paired real--simulator observations. Across controlled synthetic and physical real-world benchmarks, SPIN improves real-world posterior inference, with the improvement becoming clearer as misspecification increases.
Joon Jang, Eunho Jeong, Kyu Sung Choi +1
Department of Biomedical Sciences, Seoul National University, Seoul, Republic of Korea · Department of Applied Bioengineering, Graduate School of Convergence Science and Technology, Seoul National University, Seoul, Republic of Korea · Department of Radiology, Seoul National University Hospital, Seoul, Republic of Korea +3
Diffusion models have recently emerged as powerful learners for simulation-based inference (SBI), enabling fast and accurate estimation of latent parameters from simulated and real data. Their score-based formulation offers a flexible way to learn conditional or joint distributions over parameters and observations, thereby providing a versatile solution to various modeling problems. In this tutorial review, we synthesize recent developments on diffusion models for SBI, covering design choices for training, inference, and evaluation. We highlight opportunities created by various concepts such as guidance, score composition, flow matching, consistency models, and joint modeling. Furthermore, we discuss how efficiency and statistical accuracy are affected by noise schedules, parameterizations, and samplers. Finally, we illustrate these concepts with case studies across parameter dimensionalities, simulation budgets, and model types, and outline open questions for future research.
Jonas Arruda, Niels Bracher, Ullrich Köthe +2
Bonn Center for Mathematical Life Sciences, Life & Medical Sciences Institute, University of Bonn, Germany · Center for Modeling, Simulation, & Imaging in Medicine (CeMSIM), Rensselaer Polytechnic Institute, NY, USA · Computer Vision and Learning Lab, Heidelberg University, Germany
Simulation-based inference (SBI) of latent parameters is often hindered by simulator misspecification, the mismatch between simulated and real-world observations caused by inherent modeling simplifications. RoPE, the recent state-of-the-art for robust SBI, addresses this through optimal transport between learned representations of real and simulated observations, but requires ground-truth parameter calibration pairs that are typically unavailable in the very settings where SBI is needed. What practitioners do have is unstructured side-information such as regime labels, instruction text, and policy bulletins. We propose Misspecification-Aware Simulation-Based Inference (MA-SBI), a calibration-free framework that turns this side-channel into a posterior correction. A learned corrector maps side-channel text to an observation-space shift applied before any pre-trained amortized posterior, requiring no retraining and no parameter ground-truth. Our main theorem bounds achievable bias reduction by the mutual information between misspecification and side-channel, with a non-vacuous constant that extends to all sub-Gaussian noise via Donsker-Varadhan. On hide-the-calibration benchmarks, MA-SBI with text alone matches the oracle posterior across 10 seeds and two backbones (TOST equivalence), while RoPE given more data does not. The two approaches are complementary: where misspecification is structural and recoverable from parameter pairs, RoPE dominates, as the theory predicts. A stochastic variant improves posterior-predictive log-likelihood on real COVID and OxCGRT epidemiological data, and correctly leaves the posterior unchanged on a well-specified cognitive-science corpus.
Arunkumar V, Manoranjan Gandhudi, Gangadharan G. R. +2
University College of Engineering, Anna University Tiruchirappalli, Tamil Nadu, India · Central University of Karnataka, India · National Institute of Technology Tiruchirappalli, India +1