Active Feature Acquisition for Cost-Efficient Temporal Prediction with Reduced Participant Burden
Authors: Yunni Qu, Bing Cai Kok, Whitney Ringwald, Grant King, Aidan Wright, Kathleen Gates, Junier Oliva
Organizations: Department of Computer Science, University of North Carolina at Chapel Hill · Department of Psychology and Neuroscience, University of North Carolina at Chapel Hill · School of Social Sciences, Nanyang Technological University, Singapore · Department of Psychology, University of Minnesota Twin Cities · Department of Psychology, University of Michigan
Accurate forecasting of pathological outcomes is a central problem in psychology. To do so, psychologists often collect intensive longitudinal data. However, in such studies, the desire to acquire a large number of variables for the sake of accurate prediction is often counteracted by the need to minimize participant burden. Acquiring more variables per occasion can yield better predictions, but having too many acquisitions increase the risk of non-response and attrition. Longitudinal Active Feature Acquisition (LAFA) is a principled approach to resolve this conundrum. Instead of requiring responses to every item at every acquisition occasion, LAFA produces a policy that seeks to optimally select dynamic subsets of items to be acquired at each timepoint while preserving our ability to forecast a specific outcome. However, existing LAFA methods are mostly based on Neural Networks (NN) that are difficult to interpret in practice. In this work, we introduce a tree distillation method for learning an interpretable policy from NN-based LAFA networks. We validated our method through both a simulation and an empirical EMA dataset on forecasting daily alcohol consumption. In both cases, we find that we can meaningfully reduce the number of items acquired at each occasion with minimal loss in accuracy. Networks (NN) that are difficult to interpret in practice. In this work, we introduce a tree distillation method for learning an interpretable policy from NN-based LAFA networks. We validated our method through both a simulation and an empirical EMA dataset on forecasting daily alcohol consumption. In both cases, we find that we can meaningfully reduce the number of items acquired at each occasion with minimal loss in accuracy.
Figures & tables
Terminology
Description
Policy
A protocol that states what should be done given the existing information (e.g. if A happens, then do B; otherwise, do C).
Acquisition
Collection of data for certain select variables (e.g. at t=3 , ask the participant about their sleep and mood only).
Distillation
The process of using a simpler (and usually more interpretable) model to approximate the behavior of another model.
Teacher
The teacher is the original, complex model whose decisions we want a simpler model to imitate.
Student
The student is the simpler, more interpretable model trained to reproduce the teacher’s decisions.
Rollout
A rollout is the process of applying an acquisition policy sequentially to a participant, producing a series of states that contain the information acquired up to each occasion.
Table 1 : Glossary of terms used to describe the EMA problem in the TREACT framework.
Figure 1 : REACT Longitudinal planner. Given onboarding context and longitudinal measurements up to t, the planner outputs a binary acquisition mask over future feature–time pairs. The left portion corresponds to previously acquired measurements (unused); the right specifies the future acquisition plan under the cost-benefit objective. The earliest selected future time defines tnext .
Figure 2 : The tree as a protocal with a toy example. Here we have a subject with feature values and we trace through the subject’s path on the tree. The orange borders and branches are the subject’s path. Note that ‘sex’ is never acquired in this tree policy, but is included in the diagram because the data that TREACT is trained on contains the ‘sex’ variable.
Figure 3 : Execution of a distilled TREACT protocol across three consecutive decision occasions for one participant. At each occasion, the accumulated context and temporal history are passed to the corresponding tree, whose leaf returns an element of the codebook panel P to acquire for t. Panels 6 and 2 are acquisition patterns, and the mask panel on the right shows the resulting acquisition applied to the current occasion (previously acquired items in grey, the current acquisition in green, occasions not yet reached in blue). The middle occasion returns the "SKIP" action, so no items are administered, no cost is incurred, and the history passed forward is unchanged. A forecast yt is produced at every time step, including the t=4 on which the protocol waits, and is generated by the TREACT prediction model fϕ~ , which is finetuned classifer fϕ from REACT.
Figure 4 : AUROC/AUPRC vs cost for the simulation study. Purple dotted line indicates REACT’s classifier performance when given all of the features, acting as a empirical upper bound. The minimum cost for all the features needed to make accurate prediction on all time points is 17 for this simulation
Figure 5 : Trees of depth of 2 distilled from the REACT policy closest to the optimal minimum cost in this simulartion. Each panel P is a unique set of features. Each leaf specifies the set of features to acquire at time t , conditional on the decision rules along the path leading to that leaf.
Figure 6 : Acquisition heatmap of TREACT abd REACT with the simulation by group averaged over 5 independent runs. Yellow box indicates the ground truth to acquired for each group.
Figure 7 : AUROC/AUPRC vs cost of TREACT, REACT and baselines predicting acuiqsitions at unseen times. Purple dotted line is a empirical upper bound from the using neural network classifier of REACT with all features available. Shaded region indicate ±1 standard error of 5 independent runs.
Figure 8 : TREACT trees with depth of 3 representing acquisitions for the CHEEARS data. The policy’s AUROC is 0.750, AUPRC is 0.643, and average cost is 42.73.
Figure 9 : Heatmap of representative acquisition patterns on CHEEARS test set. We cluster on acquisition trajectories and plotted the centroids. Color shades indicate acquisition rate. Blue regions are where cross-cluster trajectories differs and orange cells are where the cell is always acuired for every participant across the clusters.
Active feature acquisition learns policies that sequentially acquire features to maximize information about a target variable. We study how to learn and evaluate such policies from finite offline data using prior-data fitted networks (PFNs), which are off-the-shelf models that output posterior predictive distributions without task-specific training. We show that under the imbalanced coverage of offline data, using total predictive entropy as a reward creates an epistemic bias that penalizes acquiring sparsely observed features. Specifically, this reward conflates epistemic uncertainty (arising from lack of offline data) with aleatoric uncertainty (arising from uninformative features). To address this, we target the posterior expected (aleatoric) entropy instead of the total predictive entropy output by a PFN for evaluating feature acquisitions. Empirical evaluations on synthetic and real-world datasets demonstrate that our approach consistently reduces value estimation bias and yields credible intervals with strong empirical coverage, which can translate to improved downstream policy selection.
Yuta Kobayashi, Divyam Madaan, Shalmali Joshi
Department of Biomedical Informatics, Columbia University
Active feature acquisition (AFA) asks which unobserved feature to measure next for each test instance under a budget. Greedy rules are easy to train but can overlook context features whose value is realized only through later acquisitions, while reinforcement-learning and generative approaches introduce difficult optimization or conditional-density estimation. We introduce \method, a deployable, supervised alternative that learns a separate candidate-conditioned risk-to-go function for every remaining budget. Starting from the one-step terminal classification risk, the functions are fitted backward with Bellman targets; inference greedily minimizes the learned terminal risk using only observed values, the mask, candidate identity, and remaining budget. A controlled non-myopic benchmark shows the expected mechanism: at budgets two and three, \method improves accuracy over its one-step ablation by 4.84±2.17 and 4.39±1.10 percentage points (mean ± standard error over five seeds). On Fashion-MNIST with 20 candidate pixels, it improves accuracy at every nontrivial reported budget on average, including 10.20±0.74 points at four acquisitions; its mean paired gain across budgets {2,4,8,12,16} is 3.50±0.37 points. A three-seed MiniBooNE study is mixed at small budgets but positive at 8 and 16 acquisitions, identifying a current boundary rather than supporting a universal claim. These results establish a reproducible mechanism-level case for direct Bellman risk regression and delimit the experiments still needed for state-of-the-art comparison.
Jiaorong Feng, Qian Li, Ying Li
Curtin Business School, Curtin University, Perth, Western Australia, Australia · †Present affiliation: Independent Researcher. · School of Electrical Engineering, Computing and Mathematical Sciences, Curtin University, Perth, Western Australia, Australia
Active feature acquisition (AFA) considers prediction problems in which features are costly to obtain and the learner adaptively decides which feature values to acquire for each instance and when to stop and predict. AFA can be formulated as a partially observable Markov decision process (POMDP), which naturally admits a sequential decision-making perspective. In this paper, we present non-myopic pathwise policy gradients (NM-PPG), a new AFA method built around this formulation. We introduce a continuous relaxation of the acquisition process that enables pathwise gradients through the full acquisition trajectory, avoiding the high variance of standard score-function policy gradients while allowing end-to-end optimization of a non-myopic acquisition policy. To better align training with deployment, we further develop a straight-through rollout scheme that follows hard feature acquisitions in the forward pass while backpropagating through the corresponding soft relaxation in the backward pass. We stabilize optimization with entropy regularization and staged temperature sharpening. Experiments on both synthetic and real-world datasets demonstrate that NM-PPG yields superior performance relative to state-of-the-art AFA baselines.
Linus Aronsson, Morteza Haghir Chehreghani
Department of Computer Science and Engineering · Chalmers University of Technology & University of Gothenburg · Gothenburg, Sweden