A Statistical Inference Framework for PMI Estimation and SGNS Word Embeddings
Organizations: Beijing Normal-Hong Kong Baptist University
Abstract
Pointwise Mutual Information (PMI) is a core measure of testing word association, and Skip-gram with Negative Sampling (SGNS) is essentially a method that implicitly factorizes a shifted PMI matrix. However, a systematic and well-rounded characterization of finite-sample uncertainty in PMI estimation remains absent and imperative to venture into. We provide a statistical framework for PMI estimation and its connection to SGNS. We prove consistency, asymptotic unbiasedness, and asymptotic normality of the empirical PMI estimator, derive its variance via the Delta method, and, applying stochastic approximation theory, obtain a variance decomposition for SGNS-based PMI estimation that separates data variance from optimization variance. Simulation experiments validate the Delta method approximation. Real-data experiments on the Brown Corpus (d = 100) reveal that SGNS systematically deviates from the theoretical relationship PMI + log K. The empirical relationship shows an attenuated PMI coefficient, an amplified log K effect, and a positive intercept, indicating systematic bias. Word analogy validation confirms the models are effective. The failure to validate the variance decomposition under low-dimensional conditions does not diminish its theoretical value; rather, it identifies the unbiasedness assumption as the key bottleneck and clarifies the gap between asymptotic theory and practice, providing implications for both practice and theory.
Figures & tables
| Specific words in the vocabulary | |
| Vocabulary size | |
| Total number of windows in the corpus | |
| Number of windows containing word | |
| Number of windows word and co-occurring | |
| Empirical probability of word | |
| Empirical probability of and co-occurring |
| Statistic | Value |
|---|---|
| Total word tokens | 981,716 |
| Unique word types | 40,234 |
| Number of documents | 500 |
| Mean document length | 2,322 words |
| Window size (sliding) | 5 |
| Total training windows ( ) | 979,716 |
| Frequency | PMI Type | Word Pairs |
|---|---|---|
| High | Positive | (is, it), (it, was), (had, he) |
| Near-zero | (of, the), (in, the), (and, the) | |
| Negative | (a, the), (in, of), (and, to) | |
| Medium | Positive | (do, something), (age, at) |
| Near-zero | (an, make), (some, which) | |
| Negative | (political, to), (had, new) |
| Semantic | Syntactic | Overall | |||
|---|---|---|---|---|---|
| Cor/Total | Accuracy | Cor/Total | Accuracy | ||
| 1 | 38/306 | 12.42% | 66/7878 | 0.84% | 1.27% |
| 2 | 49/306 | 16.01% | 46/7878 | 0.58% | 1.16% |
| 5 | 60/306 | 19.61% | 42/7878 | 0.53% | 1.25% |
| 10 | 41/306 | 13.40% | 61/7878 | 0.77% | 1.25% |
| 15 | 51/306 | 16.67% | 67/7878 | 0.85% | 1.44% |
| Correct | Total | -value | |
|---|---|---|---|
| 1 | 38 | 306 | |
| 2 | 49 | 306 | |
| 5 | 60 | 306 | |
| 10 | 41 | 306 | |
| 15 | 51 | 306 | |
| 20 | 56 | 306 |
| Category \K | ||||||
|---|---|---|---|---|---|---|
| High Positive | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| High Near-zero | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| High Negative | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| Medium Positive | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| Medium Near-zero | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| Medium Negative | 0.50 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| Metric | Theoretical | With Intercept | Without Intercept |
|---|---|---|---|
| A (PMI) | 1.0000 | 0.7077 | 0.8046 |
| B (log K) | 1.0000 | 1.5692 | 2.2174 |
| Intercept | 0.0000 | 1.5692 | 0 (fixed) |
| R² | 0.6365 | 0.5030 | |
| RMSE | 2.9221 | 1.3685 | 1.6002 |
| MAE | 2.5208 | 1.1265 | 1.3262 |
| Statistic | With Intercept | Without Intercept |
|---|---|---|
| Coefficient: A (PMI) | ||
| Estimate | ||
| t-statistic | ||
| p-value | ||
| Coefficient: B (log K) | ||
| Estimate | ||
| PMI log | SGNS | |||
|---|---|---|---|---|
| 1 | 0% | 31.4% | 29.4% | 7.40% |
| 5 | 0% | 39.3% | 36.0% | 7.13% |
| 15 | 0% | 7.80% | 6.37% | 5.97% |