Active feature acquisition learns policies that sequentially acquire features to maximize information about a target variable. We study how to learn and evaluate such policies from finite offline data using prior-data fitted networks (PFNs), which are off-the-shelf models that output posterior predictive distributions without task-specific training. We show that under the imbalanced coverage of offline data, using total predictive entropy as a reward creates an epistemic bias that penalizes acquiring sparsely observed features. Specifically, this reward conflates epistemic uncertainty (arising from lack of offline data) with aleatoric uncertainty (arising from uninformative features). To address this, we target the posterior expected (aleatoric) entropy instead of the total predictive entropy output by a PFN for evaluating feature acquisitions. Empirical evaluations on synthetic and real-world datasets demonstrate that our approach consistently reduces value estimation bias and yields credible intervals with strong empirical coverage, which can translate to improved downstream policy selection.
Figures & tables
Figure 1 : Data-driven evaluation of candidate acquisition policies using historical data of covariate and outcome pairs for Alzheimer’s risk prediction (CSF: cerebrospinal fluid, RBC: red blood cell count, WBC: white blood cell count).
Figure 2 : Uncertainty Decomposition for Value Estimation of AFA Policies. Top: Using pointwise PPD ( p(Y∣xwbc,z1:n) ) conflates aleatoric ( Ua ) and epistemic ( Ue ) uncertainties and leads to underestimation of information gain. Bottom: Our approach resolves this issue by drawing posterior samples ( P(1:K)∣z1:n ) to estimate the expected entropy over the posterior ( EP∣z1:n[H~(Y∣xwbc)] ), isolate aleatoric uncertainty Ua .
Figure 3 : Value bias and coverage for synthetic DGPs. Value estimates averaged over 1000 datasets sampled per simulated DGP class (mean ± standard error). The Total estimators suffer from epistemic bias leading to underestimation of the true feature acquisition value (negative signed bias), whereas the proposed BPI and Bootstrap approaches mitigate this bias.
Figure 4 : Downstream action selection performance on synthetic and real-world environments. We report pairwise accuracy for the simulated DGPs and clinical tasks (sampled across 200 and 100 datasets per setting, respectively). Within each sampled dataset, we evaluate candidate policies, where single-step policies are single features over the whole population, and multistep policies are stochastic policies over sequential steps. Pairwise accuracy represents the proportion of times the estimator correctly ranks the superior policy when comparing all possible pairs of candidates.
Figure 5 : Value bias for single and multistep acquisition policies on Sepsis task. We report the value bias across each panels for single-step acquisition, and bias across multistep sequential policies that acquire different number of panels.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
PhysioNet 2019
B : demographics
Age, Gender
B : vitals
HR, O 2 Sat, Temp, SBP, MAP, Resp
L : blood gas
HCO 3 , FiO 2 , pH, PaCO 2 , SaO 2
L : metabolic
Ca, Cl, Creatinine, Glucose, Mg, PO 4 , K
L : lactate
Lactate
L : CBC
WBC, Hgb, Platelets
Appendix
Table C.1: Panels in the ICU experiment. B = always-observed baseline; L = acquirable lab panels.
Synthetic
Tabular
ICU (PhysioNet)
Hyperparameter
known truth
benchmark
one-step
multistep
Worlds and sweep
Tasks / families
3 families
6
2
2
Datasets Per Task
500
100
50
50
n
(100,300) or (100,500)
(1000,2000)
{500,1000,1500}
{500,1000}
Eval patients
200
n/2
n/2
n/2
Appendix
Table C.2 : Hyperparameters of the four experiments. {⋅} denotes a uniform draw per DGP .
Pairwise accuracy
Regret
Missing rate
Variant
Δ
95% CI
p
Δ
95% CI
p
0 ( n=496 )
Total (bootstrap)
−0.005
[−0.010,−0.001]
0.993
+0.0000
[−0.0006,+0.0005]
0.718
Bootstrap
+0.004
[+0.000,+0.009]
0.027
−0.0001
[−0.0007,+0.0004]
0.241
Marginal (clt)
+0.007
[+0.001,+0.014]
0.006
−0.0001
[−0.0003,+0.0001]
0.016
Copula
−0.002
[−0.004,+0.000]
0.946
+0.0000
[−0.0002,+0.0002]
0.673
CLT
+0.005
[−0.000,+0.012]
0.013
−0.0002
[−0.0005,+0.0002]
0.157
Appendix
Table D.3 : Policy ranking of each variant against Total (plug-in) on the synthetic DGPs (Linear), one block per swept global missing rate: the mean paired difference (variant − plug-in; higher is better for pairwise accuracy, lower for regret) over DGPs with its 95% percentile bootstrap CI (DGPs resampled), and the one-sided Wilcoxon signed-rank p -value against the variant being no better than the baseline (higher pairwise accuracy, lower regret) (bold: p<0.05 . n = DGPs.)
Pairwise accuracy
Regret
Missing rate
Variant
Δ
95% CI
p
Δ
95% CI
p
0 ( n=499 )
Total (bootstrap)
+0.003
[−0.002,+0.007]
0.167
−0.0002
[−0.0008,+0.0004]
0.185
Bootstrap
+0.007
[+0.002,+0.013]
0.032
−0.0000
[−0.0006,+0.0006]
0.419
Marginal (clt)
−0.003
[−0.008,+0.003]
0.872
+0.0009
[+0.0002,+0.0018]
0.980
Copula
+0.004
[+0.000,+0.007]
0.008
+0.0001
[−0.0004,+0.0006]
0.371
CLT
−0.001
[−0.006,+0.004]
0.606
+0.0007
[+0.0000,+0.0015]
0.991
Appendix
Table D.4 : Policy ranking of each variant against Total (plug-in) on the synthetic DGPs (GMM), one block per swept global missing rate: the mean paired difference (variant − plug-in; higher is better for pairwise accuracy, lower for regret) over DGPs with its 95% percentile bootstrap CI (DGPs resampled), and the one-sided Wilcoxon signed-rank p -value against the variant being no better than the baseline (higher pairwise accuracy, lower regret) (bold: p<0.05 . n = DGPs.)
Pairwise accuracy
Regret
Missing rate
Variant
Δ
95% CI
p
Δ
95% CI
p
0 ( n=499 )
Total (bootstrap)
+0.002
[−0.003,+0.006]
0.122
−0.0001
[−0.0006,+0.0004]
0.143
Bootstrap
+0.001
[−0.003,+0.005]
0.213
+0.0002
[−0.0002,+0.0007]
0.185
Marginal (clt)
−0.002
[−0.006,+0.002]
0.783
−0.0001
[−0.0008,+0.0003]
0.499
Copula
+0.001
[−0.001,+0.003]
0.157
−0.0002
[−0.0006,+0.0000]
0.079
CLT
−0.001
[−0.004,+0.003]
0.630
−0.0001
[−0.0008,+0.0004]
0.240
Appendix
Table D.5 : Policy ranking of each variant against Total (plug-in) on the synthetic DGPs (BNN), one block per swept global missing rate: the mean paired difference (variant − plug-in; higher is better for pairwise accuracy, lower for regret) over DGPs with its 95% percentile bootstrap CI (DGPs resampled), and the one-sided Wilcoxon signed-rank p -value against the variant being no better than the baseline (higher pairwise accuracy, lower regret) (bold: p<0.05 . n = DGPs.)
Figure D.6 : Value Estimation Bias and Coverage on OpenML Tabular Benchmarks. The mean bias (absolute and signed), coverage, and interval width and standard error is computed over 500 random single-step acquisition policies from 100 sampled datasets for each task.
Figure D.7 : Policy ranking and selection on OpenML tabular benchmarks. For each task, 100 datasets are subsampled, and on each dataset 5 randomly generated candidate policies are evaluated with each estimator. Pairwise accuracy is the fraction of policy pairs ordered consistently with the reference (0.5 corresponds to random ordering). Regret is the reference value of the best candidate minus that of the policy selected by the estimator. Results are averaged over the 100 datasets.
Figure D.8 : Correlation between posterior spread (estimated epistemic uncertainty) and value underestimation. Each point represents a value estimate from a candidate policy. Several tabular datasets show no such correlation, implying that the plug-in error is either small, indicating adequate coverage or the plug-in displays systematic overconfidence. Alternatively, error is driven by sources that the posterior spread does not reflect, such as approximation error of the predictive model.
Figure D.9 : Single-Step and Multistep Acquisition Coverage on Sepsis. Coverage across each panel/feature group, averaged over 150 sampled evaluation datasets per task.
Figure D.10 : Single-Step and Multistep Acquisition Interval Width on Sepsis. Interval width across each panel/feature group, averaged over 150 sampled evaluation datasets per task.
Figure D.11 : Pairwise Accuracy on Sepsis. Single-step acquisitions and multistep stochastic policies
Figure D.12 : Regret on Sepsis. Single-step acquisitions and multistep stochastic policies
Figure D.13 : Spread of propensity estimates for single-step acquistiion across sample size on Sepsis.
Figure D.14 : Correlation between posterior spread (estimated epistemic uncertainty) and value underestimation. Each point represents a value estimate from a candidate policy.
Figure D.15 : Sample intervals of p for Copula-based martingale posterior.
Figure D.16 : Sample intervals of p for predictive CLT posterior.
Figure D.17 : Posterior contraction for single-step acquisition across sample size on Tabular Benchmarks.
Figure D.18 : Posterior contraction for single-step acquistiion across sample size on Sepsis.
Active feature acquisition (AFA) considers prediction problems in which features are costly to obtain and the learner adaptively decides which feature values to acquire for each instance and when to stop and predict. AFA can be formulated as a partially observable Markov decision process (POMDP), which naturally admits a sequential decision-making perspective. In this paper, we present non-myopic pathwise policy gradients (NM-PPG), a new AFA method built around this formulation. We introduce a continuous relaxation of the acquisition process that enables pathwise gradients through the full acquisition trajectory, avoiding the high variance of standard score-function policy gradients while allowing end-to-end optimization of a non-myopic acquisition policy. To better align training with deployment, we further develop a straight-through rollout scheme that follows hard feature acquisitions in the forward pass while backpropagating through the corresponding soft relaxation in the backward pass. We stabilize optimization with entropy regularization and staged temperature sharpening. Experiments on both synthetic and real-world datasets demonstrate that NM-PPG yields superior performance relative to state-of-the-art AFA baselines.
Linus Aronsson, Morteza Haghir Chehreghani
Department of Computer Science and Engineering · Chalmers University of Technology & University of Gothenburg · Gothenburg, Sweden
Prior-Fitted Networks (PFNs) amortize Bayesian prediction by meta-learning over a synthetic task prior, but their standard output is a posterior predictive distribution over noisy observations. For sequential decision-making, such as active learning and Bayesian optimization, acquisition should prioritize epistemic uncertainty about the latent signal rather than irreducible aleatoric observation noise. We show that this epistemic--aleatoric split is not identifiable in general from the posterior predictive distribution alone, even when that distribution is known exactly. We then exploit a distinctive advantage of PFNs: because the synthetic data-generating process is under our control, each task can contain an explicit latent signal and noise function, and the generator can provide query-level labels for both the noiseless target and the observation-noise variance. We use these labels to train a decoupled PFN with separate latent-signal and aleatoric heads. The observation-level predictive is induced by convolving the latent signal distribution with the learned noise model. Empirically, epistemic-only acquisition mitigates the failure mode of total-variance exploration in noisy and heteroscedastic settings. In matched comparisons, decoupled models usually improve over tuned observation-level baselines, with the clearest gains in HPO; in broader sweeps, a decoupled model obtains the best average rank in both HPO and synthetic BO.
Richard Bergna, Stefan Depeweg, José Miguel Hernández-Lobato
Algorithmic recourse methods typically assume that a predictive model has access to all features of an individual. In practice, decisions are often made with partial information, because features are costly to acquire. Active feature acquisition addresses cost-constrained prediction, but existing methods are explanation-agnostic: prior work provides explanations only after acquiring additional features, rather than using explanations to drive acquisition. This work flips that and treats algorithmic recourse and feature acquisition jointly. We use Markov Blanket theory to unify counterfactual, semifactual, and alterfactual explanations and to characterize how available recourse grows as features are acquired. Building on this framework, we propose an Explanation-Driven Feature Acquisition (EDFA) method that selects features by explanatory value per unit cost. The framework is further extended with distribution-free validity guarantees for recourse issued from partial information, which signal trustworthy, lower-cost recourse, along with a lower bound on the calibration data required to certify them. Experiments on 7 publicly available datasets with neural network-based predictive models show that EDFA acquires substantially fewer features than state-of-the-art AFA baselines while maintaining comparable accuracy and yielding more decision-relevant, actionable recourse. The implementation is available on GitHub.
Vinura Galwaduge, Jagath Samarabandu
Dept. of Electrical and Computer Engineering Western University London, ON, Canada