cs.LGOct 8, 2026

Marformer: A Transformer for Predicting Missing Data Distributions

Authors: Prabhav Singh, Xiheng Tom Wang, Haojun Shi, Jason Eisner

Organizations: School of Computing, College of Natural Sciences The University of Texas at Austin, Austin TX · Department of Computer Science Johns Hopkins University, Baltimore MD · Department of Computer Science Yale University, New Haven, CT

Abstract

Real decisions are made under incomplete information. If we observe only some of the random variables we need, we can predict the others. The \textbf{conditional marginals} over the missing variables are the key ingredient for computing Bayes risk and Value of Information (VOI), the expected gain from acquiring one more observation before deciding. We present the Marformer, a Transformer trained to directly predict conditional marginals given any set of observed values. Like BERT, which is trained to predict missing words from context, the Marformer constructs a hidden-vector representation for each distribution p(Xi)p(X_i) and iteratively refines it through attention to other distributions p(Xj)p(X_j). Unlike generative approaches, the Marformer does not model the full joint distribution, requires no domain knowledge of the data-generating process, and makes all predictions in a single forward pass. We evaluate across three synthetic domains with missing data---Bayesian networks, discretized multivariate Gaussians, and structured annotation data. The Marformer can match or outperform classical missing-data methods, even when those methods are given the true model family and prior that generated the synthetic data. We also evaluate on a real annotation dataset, where the Marformer outperforms the evaluated baselines at the largest training size. In both cases, the Marformer is substantially faster than the evaluated generative baselines.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AugMask: Training Diffusion Models on Incomplete Tabular Data via Stochastic Augmentation and Masking

    Jun 2, 2026Jungkyu Kim, Taeyoung Park, Kibok LeeLearning with Missing DataTabular Diffusion Models

  2. MatrixFormer: A Foundation Model for Matrix Completion

    Oct 5, 2026Dwaipayan Saha, Jacob Feitelberg, Kyuseong Choi +2Incomplete Data ImputationTransformer

  3. Transformers Can Learn Posterior Predictive Distributions In-Context

    May 26, 2026Gyeonghun Kang, Changwoo J. Lee, Xiang ChengBayesian InferenceTransformer Attention