Some documents that embedding-based topic models initially classify as noise later become founding members of emerging topics. At publication time, however, they appear as scattered points in embedding space and are difficult to distinguish from ordinary noise without the benefit of hindsight. We study whether such anticipatory outliers can be predicted prospectively, using only information available when a document first appears. We derive labels from the subsequent trajectories of outlier documents, distinguishing those that anticipate new topics from those that reinforce existing topics or remain isolated, and estimate label confidence through agreement across multiple embedding models. On two French news corpora, anticipatory outliers prove predictable at publication time. Under cross-validation, F1 rises from about 0.77 over the full eligible population to above 0.90 on high-consensus subsets, and remains at 0.76-0.80 under a strictly chronological evaluation. Predictive performance is driven mainly by geometric features capturing each outlier's position in embedding space.
Figures & tables
Corpus
Period
Articles
X posts
HydroNewsFr
20 Mar–8 Jun 2025
1,616
1,533
ClimateNewsFr
27 Mar–27 May 2025
2,445
1,046
Table 1: Summary of the two main corpora.
HydroNewsFr
ClimateNewsFr
k
Retained
Positive
Negative
F1± std
Retained
Positive
Negative
F1± std
1
1235
544
691
0.765 ±
0.013
1941
875
1066
0.778 ±
0.012
2
773
276
497
0.811 ±
0.056
1390
489
901
0.851 ±
0.016
3
508
150
358
0.895 ±
0.022
1005
273
732
0.895 ±
0.020
4
337
83
254
0.912 ±
0.069
766
182
584
0.915 ±
0.038
5
219
39
180
0.860 ±
0.116
570
115
455
0.952 ±
0.033
Table 2: Effect of the agreement threshold k at TA with XGBoost. Scores are cross-validated means ± standard deviations across folds. Bold marks the F1 scores of the selected k values used in the detailed analyses.
HydroNewsFr
ClimateNewsFr
Model
Precision
F1 -score
Recall
Precision
F1 -score
Recall
Baseline
0.246
0.395
1.000
0.171
0.291
1.000
Linear SVC
0.853
0.889
0.941
0.889
0.906
0.930
LogReg ℓ2
0.838
0.882
0.941
0.895
0.936
0.985
Decision Tree
0.803
0.835
0.891
0.867
0.898
0.937
Random Forest
0.867
0.885
0.917
0.905
0.935
0.969
Table 3: Article-level precision, F1 -score, and recall at TA under selected consensus settings: k=4 for HydroNewsFr and k=6 for ClimateNewsFr . Results are averaged over five article-level CV folds.
Corpus
Pos.
Neg.
F1 [95% CI]
Base
HydroNewsFr
42
211
0.762 [0.65, 0.86]
0.285
ClimateNewsFr
26
316
0.800 [0.66, 0.91]
0.141
Table 4: Forward-chaining results at TA with XGBoost and W=7 -day test windows, under the selected consensus settings. Cells report F1 with a 95% percentile bootstrap interval and the constant-positive baseline.
HydroNewsFr
ClimateNewsFr
Feature set
Precision
F1
Recall
Precision
F1
Recall
All features
0.919
0.912
0.917
0.970
0.969
0.969
Only geometry
0.907
0.900
0.905
0.970
0.969
0.969
Only text
0.322
0.272‡
0.242
0.279
0.202‡
0.165
Only social
0.707
0.349‡
0.239
0.235
0.083‡
0.212
All w/o geometry
0.406
0.385†
0.371
0.285
0.226†
0.194
Table 5: Ablations for XGBoost at TA . Results are 5-fold CV means. Symbols mark significant F1 drops at α=0.05 : † vs. all features; ‡ vs. geometry only.
Feature
Mean ∣SHAP∣
Spearman r
HydroNewsFr
ner_misc
0.0113
0.82∗∗∗
ner_total_ents
0.0097
−0.79∗∗∗
ner_person
0.0082
−0.81∗∗∗
text_subjectivity
0.0071
0.82∗∗∗
media_weighted_clustering
0.0027
0.53∗∗∗
Table 7: Main non-geometric SHAP features at TA . Significance coding: ∗∗∗p<0.001 .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Dim
Language
Access
sentence-camembert-base
768
French
Open
Solon-large-0.1
1024
French
Open
paraphrase-MiniLM-L12-v2
384
Multilingual
Open
paraphrase-mpnet-base-v2
768
Multilingual
Open
LaBSE
768
Multilingual
Open
multilingual-e5-large
1024
Multilingual
Open
Appendix
Table 8: Embedding models used in the experiments.
Family
Feature
Definition / interpretation
Geometric: topic distance
d1_nearest_centroid_pct
Within-model percentile rank of the Euclidean distance from the article embedding to the nearest non-noise topic centroid in the snapshot at TA .
Geometric: topic distance
d2_second_centroid_pct
Within-model percentile rank of the Euclidean distance to the second-nearest non-noise topic centroid at TA .
Geometric: topic distance
margin_d2_minus_d1_pct
Within-model percentile rank of the margin between the second-nearest and nearest centroid distances at TA .
Geometric: topic shape
mahal_nearest_pct
Within-model percentile rank of the minimum diagonal Mahalanobis distance to an existing topic cluster at TA .
Geometric: local density
knn_mean_k20_pct
Within-model percentile rank of the mean Euclidean distance to up to 20 nearest neighbors in the snapshot at TA , fewer when the snapshot is smaller.
Geometric: local density
knn_std_k20_pct
Within-model percentile rank of the standard deviation of distances to these same neighbors at TA .
Appendix
Table 9: Glossary of predictors used in the supervised models.
Table 10: Classifier hyperparameters used in the experiments. We use random_state=42 where applicable. For XGBoost, N+ and N− denote the positive and negative counts in the selected labeled subset for the corresponding corpus–consensus setting.
HydroNewsFr
ClimateNewsFr
Model
Precision
Recall
F1
Precision
Recall
F1
Baseline
0.246 ± 0.011
1.000 ± 0.000
0.395 ± 0.014
0.171 ± 0.017
1.000 ± 0.000
0.291 ± 0.025
Linear SVC
0.853 ± 0.095
0.941 ± 0.064
0.889 ± 0.056
0.889 ± 0.083
0.930 ± 0.063
0.906 ± 0.051
LogReg ℓ2
0.838 ± 0.073
0.941 ± 0.064
0.882 ± 0.040
0.895 ± 0.069
0.985 ± 0.031
0.936 ± 0.044
Decision Tree
0.803 ± 0.083
0.891 ± 0.115
0.835 ± 0.051
0.867 ± 0.085
0.937 ± 0.058
0.898 ± 0.058
Random Forest
0.867 ± 0.103
0.917 ± 0.080
0.885 ± 0.063
0.905 ± 0.079
0.969 ± 0.038
0.935 ± 0.057
Appendix
Table 11: Full article-level cross-validated performance at TA , reported as mean ± standard deviation across folds. Results use k=4 for HydroNewsFr and k=6 for ClimateNewsFr .
Corpus
k
Retained
Positive
Negative
Baseline
Linear SVC
LogReg
Tree
Forest
XGBoost
HydroNewsFr
1
1235
544
691
0.611
0.765
0.772
0.712
0.761
0.765
2
773
276
497
0.525
0.815
0.813
0.743
0.811
0.811
3
508
150
358
0.455
0.851
0.841
0.820
0.854
0.895
4
337
83
254
0.395
0.889
0.882
0.835
0.885
0.912
5
219
39
180
0.298
0.863
0.864
0.764
0.843
0.860
6
145
21
124
0.248
0.878
0.838
0.894
0.937
0.910
Appendix
Table 12: Classifier results at TA across symmetric consensus thresholds (k,k,0) . Values are cross-validated mean F1 scores.
Corpus
Comparison
ΔF1
p
q
HydroNewsFr
All vs. all w/o geometry
+0.526
7.19×10−4
1.83×10−3
All vs. only geometry
+0.011
0.536
0.577
All vs. all w/o social
−0.006
0.374
0.436
All vs. all w/o text
+0.006
0.374
0.436
Only geometry vs. only text
+0.628
3.66×10−4
1.17×10−3
Only geometry vs. only social
+0.551
1.39×10−4
5.74×10−4
Appendix
Table 13: Paired fold-level F1 comparisons for the selected ablation settings. Positive ΔF1 means that the first condition outperforms the second. Dashes indicate identical fold-level F1 values or degenerate zero differences, for which the paired t -test is not defined.
k
Retained
Positive
Negative
F1
Recall
1
2134
1056
1078
0.793
0.773
2
1390
605
785
0.871
0.865
3
965
387
578
0.922
0.931
4
656
240
416
0.923
0.910
5
465
163
302
0.942
0.939
6
321
107
214
0.981
0.981
Appendix
Table 14: XGBoost results on the extended corpus.
Corpus
Pos.
Neg.
W=5 d
W=7 d
W=14 d
W=21 d
Base
HydroNewsFr
42
211
0.776 [0.67, 0.86]
0.762 [0.65, 0.86]
0.780 [0.68, 0.87]
0.696 [0.60, 0.79]
0.285
ClimateNewsFr
26
316
0.821 [0.70, 0.92]
0.800 [0.66, 0.91]
0.737 [0.59, 0.85]
0.778 [0.64, 0.89]
0.141
HydroNewsFr Ext.
131
345
0.875 [0.83, 0.92]
0.885 [0.84, 0.92]
0.874 [0.83, 0.91]
0.873 [0.83, 0.91]
0.432
Appendix
Table 15: Forward-chaining results at TA , for test windows of W days. Ext. denotes the extended hydrogen corpus of Appendix D . Cells report pooled F1 with a 95% percentile bootstrap interval; Base is the constant-positive baseline. Counts are the pooled evaluated articles at W=7 d and vary slightly with W under the window-fit rule.