Generating a plausible clinical trajectory does not establish what would happen under a different treatment. We present ADMIT, a framework combining irregular multimodal representations, treatment-conditioned latent diffusion and explicit constraints on generated states or actions. We formulate its interventional target through sequential g-computation and distinguish causal assumptions from constraint satisfaction. Its admissibility mechanism translates physiological prior knowledge into explicit constraints on generated states and proposed actions. Treatment-exposure dynamics condition latent transitions, while state projection or action gating applies the constraints during rollout so that they influence subsequent trajectory generation. In our preliminary experiments, multimodal inputs improved supervised hidden-state recovery and reduced treatment-contrast error. In a simulated dosing-schedule experiment with leak-free history encoding, ADMIT predicted most of the tumor-volume change caused by redistributing a fixed total dose. An exposure input improved these predictions around a temporary dose reduction whether or not the assumed clearance rate was correct, but reduced the predicted size of a dose effect, and a deterministic recurrent baseline matched ADMIT's average predictions. Exposure projection reduced constraint violations, although enforcement remained incomplete. Semi-synthetic experiments using eICU context illustrated treatment-response generation under fixed and adaptive policies. Observational examples further characterize model treatment sensitivity. ADMIT provides a framework for testing whether complementary observations and physiological restrictions improve intervention trajectories, with representation recovery, effect accuracy and rule enforcement assessed separately.
Figures & tables
Figure 1: ADMIT framework. Aligned multimodal observations form latent states. Treatment-conditioned diffusion generates successive blocks, with optional constraints before updating the history. Modality-specific heads reconstruct inputs and decode selected targets. The experiments generate numerical outputs; the notes channel uses categorical tokens. The physiology guard approximately enforces a specified rule.
H visibility
Inputs
H MSE
Contrast RMSE
Contrast correlation
1
Multimodal
0.130±0.007
0.755±0.022
0.838±0.010
1
Target only
0.159±0.016
0.783±0.014
0.838±0.004
0
Multimodal
0.203±0.012
0.783±0.008
0.843±0.016
0
Target only
0.784±0.045
0.913±0.034
0.827±0.004
Table 1: Synthetic representation and treatment-contrast diagnostics. Mean ± sample standard deviation across three seeds. Visibility zero removes the direct contribution of H to target observations. The H readout is jointly supervised. Treatment contrasts compare always treating with never treating, pooled over 20 future steps and three standardized target channels.
Figure 2: Synthetic tumor trajectories. Four examples span severity at intervention after 20 steps. Solid lines and bands show generated means and sample standard deviations; dashed lines show paired simulator trajectories. Treatment plans never treat, always treat, or treat when tumor volume exceeds 12. The model updates the adaptive plan once per generated block; the simulator updates it every step. Black history points denote simulator states. The model encodes the full observed record, including future observations, so these examples do not establish forecasting accuracy.
Schedule-difference RMSE
Trajectory
Dose-effect
Model
Front-loaded
Break
Break, steps 5–8
RMSE
ratio
Zero-effect reference
0.465
0.207
0.197
—
0
ADMIT, no exposure
0.104±0.010
0.081±0.009
0.089±0.030
0.62±0.16
0.96±0.03
ADMIT, λ=0.60
0.096±0.003
0.061±0.002
0.064±0.014
0.56±0.16
0.89±0.03
ADMIT, λ=0.78 (true)
0.102±0.012
0.064±0.007
0.064±0.010
0.50±0.15
0.84±0.03
ADMIT, λ=0.90
0.107±0.012
0.064±0.006
0.065±0.007
0.48±0.12
0.80±0.02
Table 2: Predicting the effects of dose timing and dose size. Mean ± sample standard deviation across three training seeds, each scoring the same 300 test patients. The first three columns give the RMSE of predicted schedule differences and the fourth the RMSE of predicted tumor volume, averaged over the three schedules; lower is better for all four. Every model has higher trajectory error in seed 2, which dominates that column’s standard deviation. The dose-effect ratio compares predicted and simulated effects of sustained dosing at 0.6 versus 0.2; one is correct. λ is the retention used to compute the exposure input; the simulator uses 0.78.
Policy
Outcome RMSE
Effect RMSE
Effect correlation
Never treat
0.535
—
—
Always treat
0.601
0.488
0.679
Adaptive
0.498
0.357
0.724
Table 3: Treatment comparisons with real clinical context and simulated outcomes. Single training seed, 1,712 validation stays, 16 future hours and two standardized synthetic targets. Effects are relative to never treating after the forecast begins; the reference policy has zero effect by definition. The adaptive plan uses separate thresholds to start and stop treatment. Values describe the latest completed configuration and do not estimate variability across training runs.
Target
Policy contrast
Factual RMSE
Contrast/RMSE
Lactate
0.067
1.170
0.058
Urine output
0.120
0.466
0.256
Creatinine
0.018
0.398
0.045
Table 4: Sensitivity to treatment changes in observational clinical records. One-seed normalized results for 1,687 eICU stays. Policy contrast is the mean absolute difference between sustained-treatment and no-treatment predictions. Factual RMSE scores observed entries; the final column is a descriptive scale ratio.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Simulator
Tumor model of Section 4 with one continuous drug and the hidden comorbidity off; on the logit scale, each dose is centered at 0.9 times the previous dose with noise standard deviation 0.8; retention λ=e−1/4 ; 12 history and 12 forecast steps; block size four.
Patients
1,000 training, 200 validation and 300 test patients, fixed for all conditions and training seeds.
Stage 1
Encoder of Section 3 with time attention and a local kernel; 16-dimensional latent state; 60 epochs, batch size 128, learning rate 10−3 ; no exposure head; also supervised with the simulator’s noise-free target signal.
Transition
40 epochs, batch size 128, learning rate 10−3 , hidden size 64, 100 diffusion steps; rollout-loss weight three over four free-running steps; no condition dropout; guidance scale one.
GRU baseline
Deterministic recurrent network with hidden size 128 whose input is the latent state, dose and exposure; 200 epochs; trained on true previous states and on its own previous predictions.
Evaluation
64 ADMIT samples with sampling noise shared across schedules; simulator references average 32 runs with process noise shared across schedules.
Appendix
Table 5: Dosing-schedule experiment settings.
History change
Probe R2
Seed
Leak-free
Full record (max / mean)
Exposure error
None
λ=0.60
λ=0.78
λ=0.90
0
0
6.75 / 0.51
<3×10−7
0.962
0.969
0.971
0.973
1
0
4.98 / 0.54
<3×10−7
0.961
0.971
0.969
0.971
2
0
5.09 / 0.48
<3×10−7
0.961
0.969
0.973
0.969
Appendix
Table 6: Validity checks per training seed. History change: largest absolute change in the encoded history after all observations following the last history step are replaced. Exposure error: largest absolute difference between the model’s exposure ( λ=0.78 ) and the simulator’s drug level over factual doses and all planned schedules. Probe R2 : test-set fraction of variance in the simulator’s exposure explained by a linear regression on the history summary Rt , for each exposure input.
Comparison
Difference
Seed 0
Seed 1
Seed 2
No exposure − true λ
Front-loaded
− 0.019 [ − 0.023, − 0.014]
0.012 [0.006, 0.017]
0.014 [0.009, 0.018]
Break
0.003 [0.001, 0.005]
0.016 [0.013, 0.018]
0.033 [0.029, 0.036]
λ=0.60− true λ
Front-loaded
− 0.018 [ − 0.021, − 0.016]
0.003 [ − 0.000, 0.005]
− 0.002 [ − 0.005, 0.001]
Break
− 0.012 [ − 0.014, − 0.011]
0.002 [0.001, 0.003]
0.001 [ − 0.001, 0.003]
λ=0.90− true λ
Front-loaded
− 0.004 [ − 0.006, − 0.001]
0.005 [0.003, 0.007]
0.016 [0.013, 0.019]
Break
− 0.012 [ − 0.014, − 0.011]
− 0.001 [ − 0.003, 0.000]
0.012 [0.010, 0.013]
Appendix
Table 7: Paired patient-bootstrap differences in schedule-difference RMSE. For each training seed, the 300 test patients are resampled with replacement 4,000 times, with the same resample for both models. Entries give the first model’s RMSE minus the second’s, with a 95% percentile interval in brackets; negative values favor the first model. ADMIT and GRU rows without a stated input use the true λ=0.78 .