Pointwise Mutual Information (PMI) is a core measure of testing word association, and Skip-gram with Negative Sampling (SGNS) is essentially a method that implicitly factorizes a shifted PMI matrix. However, a systematic and well-rounded characterization of finite-sample uncertainty in PMI estimation remains absent and imperative to venture into. We provide a statistical framework for PMI estimation and its connection to SGNS. We prove consistency, asymptotic unbiasedness, and asymptotic normality of the empirical PMI estimator, derive its variance via the Delta method, and, applying stochastic approximation theory, obtain a variance decomposition for SGNS-based PMI estimation that separates data variance from optimization variance. Simulation experiments validate the Delta method approximation. Real-data experiments on the Brown Corpus (d = 100) reveal that SGNS systematically deviates from the theoretical relationship PMI + log K. The empirical relationship shows an attenuated PMI coefficient, an amplified log K effect, and a positive intercept, indicating systematic bias. Word analogy validation confirms the models are effective. The failure to validate the variance decomposition under low-dimensional conditions does not diminish its theoretical value; rather, it identifies the unbiasedness assumption as the key bottleneck and clarifies the gap between asymptotic theory and practice, providing implications for both practice and theory.
Figures & tables
x,y
Specific words in the vocabulary
V
Vocabulary size
N
Total number of windows in the corpus
nx
Number of windows containing word x
nxy
Number of windows word x and y co-occurring
px=nx/N
Empirical probability of word x
pxy=nxy/N
Empirical probability of x and y co-occurring
Table 1
Figure 1 : PMI Estimation and Bias
Figure 2 : PMI Estimation Variance
Figure 3 : Document length distribution of the Brown Corpus.
Statistic
Value
Total word tokens
981,716
Unique word types
40,234
Number of documents
500
Mean document length
2,322 words
Window size (sliding)
5
Total training windows ( N )
979,716
Table 1 : Brown Corpus Statistics
Figure 4 : Word frequency distribution (Zipf’s law). Left: rank–frequency plot on log–log scales. Right: histogram of log-transformed frequencies.
Figure 5 : Co-occurrence statistics. Left: histogram of co-occurrence counts ( nxy ). Right: scatter plot of PMI vs. log(nxy) .
Figure 6 : PMI distribution. Left: histogram of PMI values. Right: boxplot of PMI by co-occurrence frequency category.
Frequency
PMI Type
Word Pairs
High
Positive
(is, it), (it, was), (had, he)
Near-zero
(of, the), (in, the), (and, the)
Negative
(a, the), (in, of), (and, to)
Medium
Positive
(do, something), (age, at)
Near-zero
(an, make), (some, which)
Negative
(political, to), (had, new)
Table 2 : Selected Word Pairs by Category ( 3×3 Grid)
Figure 7 : Selected word pairs for SGNS experiments, highlighted against the background distribution.
K
Semantic
Syntactic
Overall
Cor/Total
Accuracy
Cor/Total
Accuracy
1
38/306
12.42%
66/7878
0.84%
1.27%
2
49/306
16.01%
46/7878
0.58%
1.16%
5
60/306
19.61%
42/7878
0.53%
1.25%
10
41/306
13.40%
61/7878
0.77%
1.25%
15
51/306
16.67%
67/7878
0.85%
1.44%
Table 3: Analogy accuracy for SGNS trained on Brown Corpus
Figure 8 : Analogy accuracy as a function of word frequency (figure adapted from [ 13 ] ).
K
Correct
Total
p -value
1
38
306
5.25×10−127
2
49
306
4.20×10−196
5
60
306
1.79×10−212
10
41
306
2.40×10−138
15
51
306
6.70×10−177
20
56
306
1.44×10−196
Table 4 : Binomial test results for semantic analogy accuracy
Category \K
1
2
5
10
15
20
High Positive
1.00
1.00
1.00
1.00
1.00
1.00
High Near-zero
1.00
1.00
1.00
1.00
1.00
1.00
High Negative
1.00
1.00
1.00
1.00
1.00
1.00
Medium Positive
1.00
1.00
1.00
1.00
1.00
1.00
Medium Near-zero
1.00
1.00
1.00
1.00
1.00
1.00
Medium Negative
0.50
1.00
1.00
1.00
1.00
1.00
Table 5: Rejection Rates of H0(1) by Category and K
Figure 9 : Bias Distribution of Theoretical Formula
Metric
Theoretical
With Intercept
Without Intercept
A (PMI)
1.0000
0.7077
0.8046
B (log K)
1.0000
1.5692
2.2174
Intercept
0.0000
1.5692
0 (fixed)
R²
−0.6572
0.6365
0.5030
RMSE
2.9221
1.3685
1.6002
MAE
2.5208
1.1265
1.3262
Table 6: Regression model comparison (transposed)
Statistic
With Intercept
Without Intercept
Coefficient: A (PMI)
Estimate
0.7077±0.1348
0.8046±0.1560
t-statistic
−2.17
−1.25
p-value
0.0320
0.2129
Coefficient: B (log K)
Estimate
1.5692±0.1145
2.2174±0.0719
Table 7 : Hypothesis tests for regression coefficients (transposed)
Figure 10 : Regression diagnostic plots. Top row: theoretical model ( θ^=PMI+logK ). Bottom row: empirical model with intercept ( θ^=0.70⋅PMI+1.57⋅logK+1.57 ). Left: actual vs. predicted. Middle: residuals vs. predicted. Right: coefficient comparison.
Estimating mutual information from text usually requires training a task-specific critic, which limits its use in low-data settings. We ask whether large language models can instead estimate pointwise mutual information zero-shot, using only prompts and elicited probabilities. We construct a benchmark from three publicly available human-annotated datasets with ground-truth PMI, and evaluate five information-theoretic prompting-based estimators. Our main method, PromptNCE, frames conditional probability estimation as a contrastive task and augments the candidate set with an explicit OTHER category. The OTHER category allows the model to assign probability mass outside the candidate set, avoiding the closed-set normalization of standard contrastive prompts. PromptNCE gives the best conditional probability estimates on all three datasets. For full PMI, we find that estimating label base rates is the primary bottleneck on two of the three datasets, with the best methods reaching Spearman correlation up to 0.78. We also present a case study in computer science education showing how these estimators can be used to score student knowledge summaries in a low-data setting. We release our code and prompts.
Juliette Woodrow, Chris Piech
Department of Computer Science Stanford University
Prediction-Powered Inference (PPI) is a popular strategy for combining gold-standard and possibly noisy pseudo-labels to perform statistical estimation. Prior work has shown an asymptotic \enquote{free lunch} for PPI++, an adaptive form of PPI, showing that the \textit{asymptotic} variance of PPI++ is always less than or equal to the variance obtained from using gold-standard labels alone. Notably, this result holds \textit{regardless of the quality of the pseudo-labels}. In this work, we demystify this result by conducting an exact finite-sample analysis of the estimation error of PPI++ on the mean estimation problem. We give a \enquote{no free lunch} result, characterizing the settings (and sample sizes) where PPI++ has provably worse estimation error than using gold-standard labels alone. Specifically, PPI++ will outperform if and only if the correlation between pseudo- and gold-standard is above a certain level that depends on the number of labeled samples (n). In some cases our results simplify considerably: For Gaussian data, for instance, the correlation must be at least 1/n−2 in order to see improvement. More broadly, by providing exact non-asymptotic expressions for the variance of PPI++ under sample splitting, we aim to empower practitioners to transparently reason about the benefits of PPI++ in specific applications. In experiments, we illustrate that our theoretical findings hold on real-world datasets.
Pranav Mani, Peng Xu, Zachary C. Lipton +1
Abridge AI, San Francisco, CA, USA · Machine Learning Department, Carnegie Mellon University, Pittsburgh, PA, USA · Department of Computer Science, Johns Hopkins University, Baltimore, MD, USA
Measuring statistical dependency between high-dimensional random variables is a fundamental task in data science and machine learning. Neural mutual information (MI) estimators offer a promising avenue, but they typically require costly iterative optimization for each new dataset, making them impractical for real-time applications. We present InfoAtlas, a foundation model-like architecture that eliminates this bottleneck by directly inferring MI in a single forward pass. Pretrained on large-scale synthetic data with rich dependence patterns, InfoAtlas learns to identify diverse dependence structures and predict MI directly from the dataset. Comprehensive experiments demonstrate that InfoAtlas matches state-of-the-art neural estimators in accuracy while achieving 100× speedup, can flexibly handle varying dimensions and sample sizes through a single unified model, and generalizes effectively to complex, real-world scenarios. By reformulating MI estimation as an inference task, InfoAtlas establishes a foundation for real-time dependency analysis.
Zhengyang Hu, Yanzhi Chen, Hanxiang Ren +5
The University of Hong Kong · University of Cambridge · Microsoft +2