Augmenting small concurrent studies with external or historical cohorts is attractive in drug development, where enrollment is slow, follow-up is expensive, and closely related trial or real-world data are often already available. Bayesian dynamic borrowing (BDB) provides a principled framework for adaptively controlling the influence of external data, but classical implementations often depend on hand-specified priors and MCMC-based inference, which can be computationally expensive and not generalizable. In this work, we study amortized neural posterior estimation (NPE) as a flexible alternative. A single network is pretrained on simulated current/external dataset pairs spanning covariate shift, outcome drift, and joint non-exchangeability, and then returns an approximate posterior for a scalar current-study target in a single forward pass. Through simulation studies, we find that NPE is most useful under outcome drift and joint mismatch: in the harder outcome-drift regimes, it gives up to about five-fold lower absolute bias than the best classical baseline and keeps Type I error close to nominal. After pretraining, posterior summaries are obtained in about 8 ms per dataset, roughly 103× faster than MCMC-based borrowing baselines in our timing experiment. We further analyze Alzheimer's Disease Neuroimaging Initiative (ADNI) data and show that, when mild cognitive impairment outcomes differ across cohorts, the NPE formulation recovers the later-cohort risk level in this example without claiming greater precision. Code is available at https://github.com/ChinHungScott/NPE-for-Bayesian-Dynamic-Borrowing-MLHC-.
Figures & tables
Figure 1: Proposed NPE workflow. Simulation scenarios generate current/external pairs for pretraining; at inference, the same summaries from a new pair are passed through the pretrained NPE to return a posterior for the current-study target.
Scenario
Covariates
Outcomes
SC1
exchangeable
exchangeable
SC2
partial shift
exchangeable
SC3
exchangeable
partial drift
SC4
strong shift
exchangeable
SC5
exchangeable
strong drift
SC6
strong shift
strong drift
Table 1: Two-axis taxonomy of source-mismatch regimes. SC1 is the exchangeable baseline; SC2 and SC4 isolate covariate shift at increasing strength; SC3 and SC5 isolate outcome drift at increasing strength; and SC6 combines strong covariate shift with strong outcome drift.
Scenario
Method
Bias
Bias SE
Cov.
Cov. SE
Type I
Type I SE
SC3
PSPower
0.96
0.02
0.00
0.00
1.00
0.00
SC3
IW
0.66
0.02
0.00
0.00
1.00
0.00
SC3
Commensurate
0.86
0.02
0.00
0.00
1.00
0.00
SC3
NPE
0.15
0.02
1.00
0.00
0.02
0.02
SC5
PSPower
0.95
0.01
0.00
0.00
1.00
0.00
SC5
IW
0.66
0.02
0.00
0.00
1.00
0.00
Table 2: Representative hard-setting summary ( Nc=100 ). Bias and coverage are from θ=1 ; Type I error is from θ=0 .
Figure 2: Non-null estimation error by scenario. Each spoke is one simulation scenario; values closer to the center indicate smaller error.
Figure 3: Time to a high-quality joint posterior. Per-dataset wall-clock (log scale) required to reach min(ESS(θ),ESS(ϕ))≥1000 for NPP and Commensurate; NPE is timed on its single amortized forward pass. Boxes span the IQR across B=5 replications × four distribution-shift scenarios. NPE is ∼103× faster than either MCMC sampler and uniformly so across regimes.
Cohort
Number of subjects
Conversion rate
Concurrent (ADNIGO + ADNI2)
376
0.1702
External (ADNI1)
337
0.3769
Table 3: ADNI cohort summary for the real-data borrowing example.
Method
AUC
Accuracy
Brier
LogLoss
MeanPred
MeanObs
ConcurrentOnly
0.918
0.881
0.084
0.271
0.183
0.168
ExternalOnly
0.923
0.887
0.088
0.303
0.248
0.168
Pooled
0.923
0.883
0.083
0.276
0.218
0.168
Table 4: Baseline benchmark on a held-out concurrent test set.
Quantity
Value
Posterior mean of θ
−1.5990
Posterior SD of θ
0.8585
95% interval for θ
(−3.3009,0.1082)
Observed concurrent event rate
0.1702
Observed external event rate
0.3769
Empirical concurrent logit
−1.5841
Table 5: Main ADNI real-data result for the NPE-based formulation.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Nc
Scenario
Method
Bias
RMSE
Coverage
Power
50
SC1
PSPower
0.0183
0.1268
0.90
1.00
SC1
IW
0.0365
0.1309
0.94
1.00
SC1
Commensurate
0.0183
0.1234
0.90
1.00
SC1
NPE
0.2201
0.3049
0.94
1.00
SC2
PSPower
-0.0126
0.1485
0.94
1.00
SC2
IW
-0.0223
0.1505
0.98
1.00
Appendix
Table 6: Full non-null simulation results ( θ=1 ).
Scenario
Method
Bias
RMSE
Coverage
Type I
Nc=50
SC1
PSPower
0.0187
0.1269
0.90
0.10
SC1
IW
0.0373
0.1311
0.94
0.06
SC1
Commensurate
0.0187
0.1235
0.90
0.10
SC1
NPE
0.3172
0.3811
0.90
0.10
SC2
PSPower
-0.0117
0.1485
0.94
0.06
Appendix
Table 7: Full null simulation results ( θ=0 ).
Figure 4: Additional non-null operating characteristics. Each spoke is one simulation scenario; the dashed polygon marks nominal 95% coverage when applicable.
Figure 5: Null operating characteristics by scenario. Dashed polygons indicate the nominal benchmark for each metric: 0.95 for coverage and 0.05 for Type I error.
Noise SD
Bias
Absolute bias
RMSE
Coverage
0.00
0.222
0.222
0.273
0.98
0.02
0.226
0.226
0.276
0.98
0.05
0.241
0.241
0.290
0.98
0.10
0.270
0.270
0.328
0.98
0.15
0.296
0.296
0.357
0.96
Appendix
Table 8: Sensitivity to perturbing the OLS-derived coefficient-difference summary block in SC5 ( nc=100 , θ=1 ).
Scenario
Method
Bias
RMSE
Coverage
Width
SC1
Concurrent OLS
-0.006
0.134
0.94
0.538
SC1
NPE- β1
0.002
0.082
1.00
0.442
SC2
Concurrent OLS
0.003
0.136
0.92
0.542
SC2
NPE- β1
0.016
0.064
1.00
0.459
SC3
Concurrent OLS
-0.002
0.142
0.94
0.530
SC3
NPE- β1
0.004
0.069
0.98
0.434
Appendix
Table 9: Limited illustrative check for a current-study covariate-effect target, β1 ( nc=100 , ne=200 ). NPE uses target-aware summaries and validation-based posterior-scale calibration.
sd(δ)
Posterior mean
Mean SD
95% interval
Width
0.50
-1.573
0.864
(−3.281,0.122)
3.403
0.75
-1.583
0.862
(−3.287,0.107)
3.394
1.00
-1.587
0.856
(−3.279,0.092)
3.371
1.25
-1.592
0.854
(−3.279,0.082)
3.360
Appendix
Table 10: ADNI sensitivity to the prior standard deviation for the external–concurrent outcome-shift parameter δ .
Scenario
Method
Bias
RMSE
Power
SC3
NPE
0.1320
0.2223
1.00
SC3
KL-NPE
0.0242
0.1905
0.98
SC3
OT-NPE
0.0757
0.2013
1.00
SC3
KL+KS-NPE
0.0964
0.2037
1.00
SC3
JSD-NPE
0.1769
0.2498
1.00
SC5
NPE
0.1483
0.2205
1.00
Appendix
Table 11: Exploratory NPE extensions in the harder non-null settings ( Nc=100 , θ=1 ).
Department of Statistics and Data Science, Indian Institute of Technology Kanpur, Kanpur, India 208016 · Marshall School of Business, University of Southern California, Los Angeles, United States 90089-0809
Department of Statistical Science, University College London, UK · The Alan Turing Institute, UK · Department of Computer Science, Aalto University, Finland