stat.MLMay 13, 2026

LLMs as Implicit Imputers: Uncertainty Should Scale with Missing Information

Authors: Stef van Buuren

Organizations: 1. TNO - Netherlands Organization for Applied Scientific Research, Leiden · 2. Dept. of Methodology and Statistics, University of Utrecht

Abstract

Large language models (LLMs) are increasingly deployed in settings where the available context is incomplete or degraded. We argue that an LLM generating answers under incomplete context can be viewed as an implicit imputer, and evaluated against a criterion from the multiple imputation (MI) literature: uncertainty should scale with the amount of missing information. We assess this criterion on SQuAD, using a controlled framework in which context availability is varied across five levels. We evaluate two answer-level uncertainty measures that can be estimated from repeated sampling: sampling-based confidence (empirical mode frequency) and response entropy. Confidence fails to reflect increasing missingness: it remains high even as accuracy collapses. Entropy, by contrast, increases with context removal, consistent with the MI analogy, and explains substantially more variance in accuracy than confidence across all evidence levels (quadratic R2R^2 gap up to 0.057). We further introduce a black-box diagnostic ρR(α)ρ_R(α) that estimates the proportion of baseline uncertainty resolved by context level αα, requiring only repeated sampling with and without context. These results suggest that entropy is a more responsive black-box uncertainty measure than confidence under incomplete context.

Explore similar work

CardsList