stat.MLOct 4, 2026

A Statistical Inference Framework for PMI Estimation and SGNS Word Embeddings

Authors: Zhongqi Fan

Organizations: Beijing Normal-Hong Kong Baptist University

Abstract

Pointwise Mutual Information (PMI) is a core measure of testing word association, and Skip-gram with Negative Sampling (SGNS) is essentially a method that implicitly factorizes a shifted PMI matrix. However, a systematic and well-rounded characterization of finite-sample uncertainty in PMI estimation remains absent and imperative to venture into. We provide a statistical framework for PMI estimation and its connection to SGNS. We prove consistency, asymptotic unbiasedness, and asymptotic normality of the empirical PMI estimator, derive its variance via the Delta method, and, applying stochastic approximation theory, obtain a variance decomposition for SGNS-based PMI estimation that separates data variance from optimization variance. Simulation experiments validate the Delta method approximation. Real-data experiments on the Brown Corpus (d = 100) reveal that SGNS systematically deviates from the theoretical relationship PMI + log K. The empirical relationship shows an attenuated PMI coefficient, an amplified log K effect, and a positive intercept, indicating systematic bias. Word analogy validation confirms the models are effective. The failure to validate the variance decomposition under low-dimensional conditions does not diminish its theoretical value; rather, it identifies the unbiasedness assumption as the key bottleneck and clarifies the gap between asymptotic theory and practice, providing implications for both practice and theory.

Figures & tables

Explore similar work

May 20, 2026cs.CL

PromptNCE: Conditional Probabilities and PMI Using Only LLMs and Contrastive Estimation Prompts

Estimating mutual information from text usually requires training a task-specific critic, which limits its use in low-data settings. We ask whether large language models can instead estimate pointwise mutual information zero-shot, using only prompts and elicited probabilities. We construct a benchmark from three publicly available human-annotated datasets with ground-truth PMI, and evaluate five information-theoretic prompting-based estimators. Our main method, PromptNCE, frames conditional probability estimation as a contrastive task and augments the candidate set with an explicit OTHER category. The OTHER category allows the model to assign probability mass outside the candidate set, avoiding the closed-set normalization of standard contrastive prompts. PromptNCE gives the best conditional probability estimates on all three datasets. For full PMI, we find that estimating label base rates is the primary bottleneck on two of the three datasets, with the best methods reaching Spearman correlation up to 0.78. We also present a case study in computer science education showing how these estimators can be used to score student knowledge summaries in a low-data setting. We release our code and prompts.
May 26, 2025stat.ML

No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered Inference

Prediction-Powered Inference (PPI) is a popular strategy for combining gold-standard and possibly noisy pseudo-labels to perform statistical estimation. Prior work has shown an asymptotic \enquote{free lunch} for PPI++, an adaptive form of PPI, showing that the \textit{asymptotic} variance of PPI++ is always less than or equal to the variance obtained from using gold-standard labels alone. Notably, this result holds \textit{regardless of the quality of the pseudo-labels}. In this work, we demystify this result by conducting an exact finite-sample analysis of the estimation error of PPI++ on the mean estimation problem. We give a \enquote{no free lunch} result, characterizing the settings (and sample sizes) where PPI++ has provably worse estimation error than using gold-standard labels alone. Specifically, PPI++ will outperform if and only if the correlation between pseudo- and gold-standard is above a certain level that depends on the number of labeled samples (nn). In some cases our results simplify considerably: For Gaussian data, for instance, the correlation must be at least 1/n−21/\sqrt{n - 2} in order to see improvement. More broadly, by providing exact non-asymptotic expressions for the variance of PPI++ under sample splitting, we aim to empower practitioners to transparently reason about the benefits of PPI++ in specific applications. In experiments, we illustrate that our theoretical findings hold on real-world datasets.
May 29, 2026cs.LG

InfoAtlas: A Foundation Model for Zero-Shot Statistical Dependence Estimate

Measuring statistical dependency between high-dimensional random variables is a fundamental task in data science and machine learning. Neural mutual information (MI) estimators offer a promising avenue, but they typically require costly iterative optimization for each new dataset, making them impractical for real-time applications. We present InfoAtlas, a foundation model-like architecture that eliminates this bottleneck by directly inferring MI in a single forward pass. Pretrained on large-scale synthetic data with rich dependence patterns, InfoAtlas learns to identify diverse dependence structures and predict MI directly from the dataset. Comprehensive experiments demonstrate that InfoAtlas matches state-of-the-art neural estimators in accuracy while achieving 100×100\times speedup, can flexibly handle varying dimensions and sample sizes through a single unified model, and generalizes effectively to complex, real-world scenarios. By reformulating MI estimation as an inference task, InfoAtlas establishes a foundation for real-time dependency analysis.