Active Feature Acquisition for Cost-Efficient Temporal Prediction with Reduced Participant Burden
Organizations: Department of Computer Science, University of North Carolina at Chapel Hill · Department of Psychology and Neuroscience, University of North Carolina at Chapel Hill · School of Social Sciences, Nanyang Technological University, Singapore · Department of Psychology, University of Minnesota Twin Cities · Department of Psychology, University of Michigan
Abstract
Accurate forecasting of pathological outcomes is a central problem in psychology. To do so, psychologists often collect intensive longitudinal data. However, in such studies, the desire to acquire a large number of variables for the sake of accurate prediction is often counteracted by the need to minimize participant burden. Acquiring more variables per occasion can yield better predictions, but having too many acquisitions increase the risk of non-response and attrition. Longitudinal Active Feature Acquisition (LAFA) is a principled approach to resolve this conundrum. Instead of requiring responses to every item at every acquisition occasion, LAFA produces a policy that seeks to optimally select dynamic subsets of items to be acquired at each timepoint while preserving our ability to forecast a specific outcome. However, existing LAFA methods are mostly based on Neural Networks (NN) that are difficult to interpret in practice. In this work, we introduce a tree distillation method for learning an interpretable policy from NN-based LAFA networks. We validated our method through both a simulation and an empirical EMA dataset on forecasting daily alcohol consumption. In both cases, we find that we can meaningfully reduce the number of items acquired at each occasion with minimal loss in accuracy. Networks (NN) that are difficult to interpret in practice. In this work, we introduce a tree distillation method for learning an interpretable policy from NN-based LAFA networks. We validated our method through both a simulation and an empirical EMA dataset on forecasting daily alcohol consumption. In both cases, we find that we can meaningfully reduce the number of items acquired at each occasion with minimal loss in accuracy.
Figures & tables
| Terminology | Description |
|---|---|
| Policy | A protocol that states what should be done given the existing information (e.g. if A happens, then do B; otherwise, do C). |
| Acquisition | Collection of data for certain select variables (e.g. at , ask the participant about their sleep and mood only). |
| Distillation | The process of using a simpler (and usually more interpretable) model to approximate the behavior of another model. |
| Teacher | The teacher is the original, complex model whose decisions we want a simpler model to imitate. |
| Student | The student is the simpler, more interpretable model trained to reproduce the teacher’s decisions. |
| Rollout | A rollout is the process of applying an acquisition policy sequentially to a participant, producing a series of states that contain the information acquired up to each occasion. |