Active feature acquisition learns policies that sequentially acquire features to maximize information about a target variable. We study how to learn and evaluate such policies from finite offline data using prior-data fitted networks (PFNs), which are off-the-shelf models that output posterior predictive distributions without task-specific training. We show that under the imbalanced coverage of offline data, using total predictive entropy as a reward creates an epistemic bias that penalizes acquiring sparsely observed features. Specifically, this reward conflates epistemic uncertainty (arising from lack of offline data) with aleatoric uncertainty (arising from uninformative features). To address this, we target the posterior expected (aleatoric) entropy instead of the total predictive entropy output by a PFN for evaluating feature acquisitions. Empirical evaluations on synthetic and real-world datasets demonstrate that our approach consistently reduces value estimation bias and yields credible intervals with strong empirical coverage, which can translate to improved downstream policy selection.
Figures & tables
Figure 1 : Data-driven evaluation of candidate acquisition policies using historical data of covariate and outcome pairs for Alzheimer’s risk prediction (CSF: cerebrospinal fluid, RBC: red blood cell count, WBC: white blood cell count).
Figure 2 : Uncertainty Decomposition for Value Estimation of AFA Policies. Top: Using pointwise PPD ( p(Y∣xwbc,z1:n) ) conflates aleatoric ( Ua ) and epistemic ( Ue ) uncertainties and leads to underestimation of information gain. Bottom: Our approach resolves this issue by drawing posterior samples ( P(1:K)∣z1:n ) to estimate the expected entropy over the posterior ( EP∣z1:n[H~(Y∣xwbc)] ), isolate aleatoric uncertainty Ua .
Figure 3 : Value bias and coverage for synthetic DGPs. Value estimates averaged over 1000 datasets sampled per simulated DGP class (mean ± standard error). The Total estimators suffer from epistemic bias leading to underestimation of the true feature acquisition value (negative signed bias), whereas the proposed BPI and Bootstrap approaches mitigate this bias.
Figure 4 : Downstream action selection performance on synthetic and real-world environments. We report pairwise accuracy for the simulated DGPs and clinical tasks (sampled across 200 and 100 datasets per setting, respectively). Within each sampled dataset, we evaluate candidate policies, where single-step policies are single features over the whole population, and multistep policies are stochastic policies over sequential steps. Pairwise accuracy represents the proportion of times the estimator correctly ranks the superior policy when comparing all possible pairs of candidates.
Figure 5 : Value bias for single and multistep acquisition policies on Sepsis task. We report the value bias across each panels for single-step acquisition, and bias across multistep sequential policies that acquire different number of panels.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
PhysioNet 2019
B : demographics
Age, Gender
B : vitals
HR, O 2 Sat, Temp, SBP, MAP, Resp
L : blood gas
HCO 3 , FiO 2 , pH, PaCO 2 , SaO 2
L : metabolic
Ca, Cl, Creatinine, Glucose, Mg, PO 4 , K
L : lactate
Lactate
L : CBC
WBC, Hgb, Platelets
Appendix
Table C.1: Panels in the ICU experiment. B = always-observed baseline; L = acquirable lab panels.
Synthetic
Tabular
ICU (PhysioNet)
Hyperparameter
known truth
benchmark
one-step
multistep
Worlds and sweep
Tasks / families
3 families
6
2
2
Datasets Per Task
500
100
50
50
n
(100,300) or (100,500)
(1000,2000)
{500,1000,1500}
{500,1000}
Eval patients
200
n/2
n/2
n/2
Appendix
Table C.2 : Hyperparameters of the four experiments. {⋅} denotes a uniform draw per DGP .
Pairwise accuracy
Regret
Missing rate
Variant
Δ
95% CI
p
Δ
95% CI
p
0 ( n=496 )
Total (bootstrap)
−0.005
[−0.010,−0.001]
0.993
+0.0000
[−0.0006,+0.0005]
0.718
Bootstrap
+0.004
[+0.000,+0.009]
0.027
−0.0001
[−0.0007,+0.0004]
0.241
Marginal (clt)
+0.007
[+0.001,+0.014]
0.006
−0.0001
[−0.0003,+0.0001]
0.016
Copula
−0.002
[−0.004,+0.000]
0.946
+0.0000
[−0.0002,+0.0002]
0.673
CLT
+0.005
[−0.000,+0.012]
0.013
−0.0002
[−0.0005,+0.0002]
0.157
Appendix
Table D.3 : Policy ranking of each variant against Total (plug-in) on the synthetic DGPs (Linear), one block per swept global missing rate: the mean paired difference (variant − plug-in; higher is better for pairwise accuracy, lower for regret) over DGPs with its 95% percentile bootstrap CI (DGPs resampled), and the one-sided Wilcoxon signed-rank p -value against the variant being no better than the baseline (higher pairwise accuracy, lower regret) (bold: p<0.05 . n = DGPs.)
Pairwise accuracy
Regret
Missing rate
Variant
Δ
95% CI
p
Δ
95% CI
p
0 ( n=499 )
Total (bootstrap)
+0.003
[−0.002,+0.007]
0.167
−0.0002
[−0.0008,+0.0004]
0.185
Bootstrap
+0.007
[+0.002,+0.013]
0.032
−0.0000
[−0.0006,+0.0006]
0.419
Marginal (clt)
−0.003
[−0.008,+0.003]
0.872
+0.0009
[+0.0002,+0.0018]
0.980
Copula
+0.004
[+0.000,+0.007]
0.008
+0.0001
[−0.0004,+0.0006]
0.371
CLT
−0.001
[−0.006,+0.004]
0.606
+0.0007
[+0.0000,+0.0015]
0.991
Appendix
Table D.4 : Policy ranking of each variant against Total (plug-in) on the synthetic DGPs (GMM), one block per swept global missing rate: the mean paired difference (variant − plug-in; higher is better for pairwise accuracy, lower for regret) over DGPs with its 95% percentile bootstrap CI (DGPs resampled), and the one-sided Wilcoxon signed-rank p -value against the variant being no better than the baseline (higher pairwise accuracy, lower regret) (bold: p<0.05 . n = DGPs.)
Pairwise accuracy
Regret
Missing rate
Variant
Δ
95% CI
p
Δ
95% CI
p
0 ( n=499 )
Total (bootstrap)
+0.002
[−0.003,+0.006]
0.122
−0.0001
[−0.0006,+0.0004]
0.143
Bootstrap
+0.001
[−0.003,+0.005]
0.213
+0.0002
[−0.0002,+0.0007]
0.185
Marginal (clt)
−0.002
[−0.006,+0.002]
0.783
−0.0001
[−0.0008,+0.0003]
0.499
Copula
+0.001
[−0.001,+0.003]
0.157
−0.0002
[−0.0006,+0.0000]
0.079
CLT
−0.001
[−0.004,+0.003]
0.630
−0.0001
[−0.0008,+0.0004]
0.240
Appendix
Table D.5 : Policy ranking of each variant against Total (plug-in) on the synthetic DGPs (BNN), one block per swept global missing rate: the mean paired difference (variant − plug-in; higher is better for pairwise accuracy, lower for regret) over DGPs with its 95% percentile bootstrap CI (DGPs resampled), and the one-sided Wilcoxon signed-rank p -value against the variant being no better than the baseline (higher pairwise accuracy, lower regret) (bold: p<0.05 . n = DGPs.)
Figure D.6 : Value Estimation Bias and Coverage on OpenML Tabular Benchmarks. The mean bias (absolute and signed), coverage, and interval width and standard error is computed over 500 random single-step acquisition policies from 100 sampled datasets for each task.
Figure D.7 : Policy ranking and selection on OpenML tabular benchmarks. For each task, 100 datasets are subsampled, and on each dataset 5 randomly generated candidate policies are evaluated with each estimator. Pairwise accuracy is the fraction of policy pairs ordered consistently with the reference (0.5 corresponds to random ordering). Regret is the reference value of the best candidate minus that of the policy selected by the estimator. Results are averaged over the 100 datasets.
Figure D.8 : Correlation between posterior spread (estimated epistemic uncertainty) and value underestimation. Each point represents a value estimate from a candidate policy. Several tabular datasets show no such correlation, implying that the plug-in error is either small, indicating adequate coverage or the plug-in displays systematic overconfidence. Alternatively, error is driven by sources that the posterior spread does not reflect, such as approximation error of the predictive model.
Figure D.9 : Single-Step and Multistep Acquisition Coverage on Sepsis. Coverage across each panel/feature group, averaged over 150 sampled evaluation datasets per task.
Figure D.10 : Single-Step and Multistep Acquisition Interval Width on Sepsis. Interval width across each panel/feature group, averaged over 150 sampled evaluation datasets per task.
Figure D.11 : Pairwise Accuracy on Sepsis. Single-step acquisitions and multistep stochastic policies
Figure D.12 : Regret on Sepsis. Single-step acquisitions and multistep stochastic policies
Figure D.13 : Spread of propensity estimates for single-step acquistiion across sample size on Sepsis.
Figure D.14 : Correlation between posterior spread (estimated epistemic uncertainty) and value underestimation. Each point represents a value estimate from a candidate policy.
Figure D.15 : Sample intervals of p for Copula-based martingale posterior.
Figure D.16 : Sample intervals of p for predictive CLT posterior.
Figure D.17 : Posterior contraction for single-step acquisition across sample size on Tabular Benchmarks.
Figure D.18 : Posterior contraction for single-step acquistiion across sample size on Sepsis.