OrthoGen: A Generative Orthogonal Learner for Time-Varying Treatments
Organizations: Novartis · Barcelona Supercomputing Center · Universitat Polit`ecnica de Catalunya · LMU Munich · Munich Center for Machine Learning
Abstract
Estimating conditional distributional potential outcomes (CDPOs) over time is important in medicine (e.g., to estimate patient-specific risks under different treatment sequences). However, this task is challenging because of time-varying confounding, yet existing adjustment strategies for this task are limited. In this paper, we aim to learn CDPOs under time-varying treatments using flexible generative models. Our contributions are two-fold. (1) We introduce a tailored adjustment strategy for our setting, namely, generative recursive g-computation. Our adjustment strategy recursively propagates full conditional outcome distributions rather than conditional means, modeling the variables of interest directly rather than full trajectories. Building on our adjustment strategy, we formulate simple generative learners for CDPO estimation. However, these learners can be sensitive to nuisance estimation errors, which motivates an orthogonal learner. (2) We thus introduce OrthoGen, a Neyman-orthogonal and doubly robust generative learner. Importantly, we show that OrthoGen further achieves rate double robustness and quasi-oracle efficiency under suitable conditions. Our learners are flexible and can be instantiated with different generative backbones (e.g., normalizing flows and diffusion models). Across experiments with synthetic, semi-synthetic and real-world datasets, we find that OrthoGen is highly effective. To the best of our knowledge, we are the first to propose a generative orthogonal learner for estimating CDPOs under time-varying treatments.
Figures & tables
| Method | Adjustment | Backbone | Modeling target | Neyman-orthogonality |
| Wu et al. (2024) | IPTW | CVAE / DM | ✗ Marginal DPO | ✗ |
| Mu et al. (2025) | IPTW | DM | ✓ CDPO | ✗ |
| PI ( ours ) | G.R. g-comp. | Model agnostic | ✓ CDPO | ✗ |
| RA ( ours ) | G.R. g-comp. | Model agnostic | ✓ CDPO | ✗ |
| OrthoGen ( ours ) | DR | Model agnostic | ✓ CDPO | ✓ |
| Notes: G.R. g-comp. = generative recursive g-computation; CVAE = conditional variational autoencoder; DM = diffusion model; DPO = distributional potential outcome; PI = plug-in; RA = regression-adjusted; DR = doubly robust. | ||||
| Method | ||||||
|---|---|---|---|---|---|---|
| NFs | PI | |||||
| RA | ||||||
| IPTW | ||||||
| OrthoGen | ||||||
| DMs | PI | |||||
| RA |
| Method | ||||||
|---|---|---|---|---|---|---|
| NFs | PI | |||||
| RA | ||||||
| IPTW | ||||||
| OrthoGen | ||||||
| DMs | PI | |||||
| RA |
| Method | ||||
|---|---|---|---|---|
| NFs | PI | |||
| RA | ||||
| IPTW | ||||
| OrthoGen | ||||
| DMs | PI | |||
| RA |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Hyperparameter | Value | Role |
|---|---|---|
| Transformer blocks | Causal encoder block | |
| Model dimension | Token and attention representation width | |
| Attention heads | Head width | |
| Feed-forward dimension | Hidden width inside the encoder block | |
| Dropout probability | Applied to the input embedding and encoder block | |
| Readout dimension | Dimension of the context passed to the downstream head |
| Family | Hyperparameter | Values considered |
| NF | Conditioner width | |
| Spline bins | ||
| Learning rate | ||
| Batch size | ||
| Weight decay | ||
| Context noise |
| Dataset | Family | Selected architecture | Selected optimization and regularization |
|---|---|---|---|
| Semi-synthetic MIMIC-III | NF | Width ; bins | Learning rate ; batch ; weight decay ; epochs; |
| DM | ; denoiser width | Learning rate ; weight decay ; | |
| Real-world MIMIC-III | NF | Width ; bins | Learning rate ; batch ; weight decay ; epochs; |
| DM | ; denoiser width | Learning rate ; weight decay ; | |
| Fully synthetic | NF | Width ; bins | Learning rate ; batch ; weight decay ; epochs; |
| DM | ; denoiser width | Learning rate ; weight decay ; |
| Dataset | NFs | DMs |
|---|---|---|
| Fully synthetic | min | min |
| Semi-synthetic MIMIC-III | min | min |
| Real-world MIMIC-III | min | min |