Some documents that embedding-based topic models initially classify as noise later become founding members of emerging topics. At publication time, however, they appear as scattered points in embedding space and are difficult to distinguish from ordinary noise without the benefit of hindsight. We study whether such anticipatory outliers can be predicted prospectively, using only information available when a document first appears. We derive labels from the subsequent trajectories of outlier documents, distinguishing those that anticipate new topics from those that reinforce existing topics or remain isolated, and estimate label confidence through agreement across multiple embedding models. On two French news corpora, anticipatory outliers prove predictable at publication time. Under cross-validation, F1 rises from about 0.77 over the full eligible population to above 0.90 on high-consensus subsets, and remains at 0.76-0.80 under a strictly chronological evaluation. Predictive performance is driven mainly by geometric features capturing each outlier's position in embedding space.
Figures & tables
Corpus
Period
Articles
X posts
HydroNewsFr
20 Mar–8 Jun 2025
1,616
1,533
ClimateNewsFr
27 Mar–27 May 2025
2,445
1,046
Table 1: Summary of the two main corpora.
HydroNewsFr
ClimateNewsFr
k
Retained
Positive
Negative
F1± std
Retained
Positive
Negative
F1± std
1
1235
544
691
0.765 ±
0.013
1941
875
1066
0.778 ±
0.012
2
773
276
497
0.811 ±
0.056
1390
489
901
0.851 ±
0.016
3
508
150
358
0.895 ±
0.022
1005
273
732
0.895 ±
0.020
4
337
83
254
0.912 ±
0.069
766
182
584
0.915 ±
0.038
5
219
39
180
0.860 ±
0.116
570
115
455
0.952 ±
0.033
Table 2: Effect of the agreement threshold k at TA with XGBoost. Scores are cross-validated means ± standard deviations across folds. Bold marks the F1 scores of the selected k values used in the detailed analyses.
HydroNewsFr
ClimateNewsFr
Model
Precision
F1 -score
Recall
Precision
F1 -score
Recall
Baseline
0.246
0.395
1.000
0.171
0.291
1.000
Linear SVC
0.853
0.889
0.941
0.889
0.906
0.930
LogReg ℓ2
0.838
0.882
0.941
0.895
0.936
0.985
Decision Tree
0.803
0.835
0.891
0.867
0.898
0.937
Random Forest
0.867
0.885
0.917
0.905
0.935
0.969
Table 3: Article-level precision, F1 -score, and recall at TA under selected consensus settings: k=4 for HydroNewsFr and k=6 for ClimateNewsFr . Results are averaged over five article-level CV folds.
Corpus
Pos.
Neg.
F1 [95% CI]
Base
HydroNewsFr
42
211
0.762 [0.65, 0.86]
0.285
ClimateNewsFr
26
316
0.800 [0.66, 0.91]
0.141
Table 4: Forward-chaining results at TA with XGBoost and W=7 -day test windows, under the selected consensus settings. Cells report F1 with a 95% percentile bootstrap interval and the constant-positive baseline.
HydroNewsFr
ClimateNewsFr
Feature set
Precision
F1
Recall
Precision
F1
Recall
All features
0.919
0.912
0.917
0.970
0.969
0.969
Only geometry
0.907
0.900
0.905
0.970
0.969
0.969
Only text
0.322
0.272‡
0.242
0.279
0.202‡
0.165
Only social
0.707
0.349‡
0.239
0.235
0.083‡
0.212
All w/o geometry
0.406
0.385†
0.371
0.285
0.226†
0.194
Table 5: Ablations for XGBoost at TA . Results are 5-fold CV means. Symbols mark significant F1 drops at α=0.05 : † vs. all features; ‡ vs. geometry only.
Feature
Mean ∣SHAP∣
Spearman r
HydroNewsFr
ner_misc
0.0113
0.82∗∗∗
ner_total_ents
0.0097
−0.79∗∗∗
ner_person
0.0082
−0.81∗∗∗
text_subjectivity
0.0071
0.82∗∗∗
media_weighted_clustering
0.0027
0.53∗∗∗
Table 7: Main non-geometric SHAP features at TA . Significance coding: ∗∗∗p<0.001 .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Dim
Language
Access
sentence-camembert-base
768
French
Open
Solon-large-0.1
1024
French
Open
paraphrase-MiniLM-L12-v2
384
Multilingual
Open
paraphrase-mpnet-base-v2
768
Multilingual
Open
LaBSE
768
Multilingual
Open
multilingual-e5-large
1024
Multilingual
Open
Appendix
Table 8: Embedding models used in the experiments.
Family
Feature
Definition / interpretation
Geometric: topic distance
d1_nearest_centroid_pct
Within-model percentile rank of the Euclidean distance from the article embedding to the nearest non-noise topic centroid in the snapshot at TA .
Geometric: topic distance
d2_second_centroid_pct
Within-model percentile rank of the Euclidean distance to the second-nearest non-noise topic centroid at TA .
Geometric: topic distance
margin_d2_minus_d1_pct
Within-model percentile rank of the margin between the second-nearest and nearest centroid distances at TA .
Geometric: topic shape
mahal_nearest_pct
Within-model percentile rank of the minimum diagonal Mahalanobis distance to an existing topic cluster at TA .
Geometric: local density
knn_mean_k20_pct
Within-model percentile rank of the mean Euclidean distance to up to 20 nearest neighbors in the snapshot at TA , fewer when the snapshot is smaller.
Geometric: local density
knn_std_k20_pct
Within-model percentile rank of the standard deviation of distances to these same neighbors at TA .
Appendix
Table 9: Glossary of predictors used in the supervised models.
Table 10: Classifier hyperparameters used in the experiments. We use random_state=42 where applicable. For XGBoost, N+ and N− denote the positive and negative counts in the selected labeled subset for the corresponding corpus–consensus setting.
HydroNewsFr
ClimateNewsFr
Model
Precision
Recall
F1
Precision
Recall
F1
Baseline
0.246 ± 0.011
1.000 ± 0.000
0.395 ± 0.014
0.171 ± 0.017
1.000 ± 0.000
0.291 ± 0.025
Linear SVC
0.853 ± 0.095
0.941 ± 0.064
0.889 ± 0.056
0.889 ± 0.083
0.930 ± 0.063
0.906 ± 0.051
LogReg ℓ2
0.838 ± 0.073
0.941 ± 0.064
0.882 ± 0.040
0.895 ± 0.069
0.985 ± 0.031
0.936 ± 0.044
Decision Tree
0.803 ± 0.083
0.891 ± 0.115
0.835 ± 0.051
0.867 ± 0.085
0.937 ± 0.058
0.898 ± 0.058
Random Forest
0.867 ± 0.103
0.917 ± 0.080
0.885 ± 0.063
0.905 ± 0.079
0.969 ± 0.038
0.935 ± 0.057
Appendix
Table 11: Full article-level cross-validated performance at TA , reported as mean ± standard deviation across folds. Results use k=4 for HydroNewsFr and k=6 for ClimateNewsFr .
Corpus
k
Retained
Positive
Negative
Baseline
Linear SVC
LogReg
Tree
Forest
XGBoost
HydroNewsFr
1
1235
544
691
0.611
0.765
0.772
0.712
0.761
0.765
2
773
276
497
0.525
0.815
0.813
0.743
0.811
0.811
3
508
150
358
0.455
0.851
0.841
0.820
0.854
0.895
4
337
83
254
0.395
0.889
0.882
0.835
0.885
0.912
5
219
39
180
0.298
0.863
0.864
0.764
0.843
0.860
6
145
21
124
0.248
0.878
0.838
0.894
0.937
0.910
Appendix
Table 12: Classifier results at TA across symmetric consensus thresholds (k,k,0) . Values are cross-validated mean F1 scores.
Corpus
Comparison
ΔF1
p
q
HydroNewsFr
All vs. all w/o geometry
+0.526
7.19×10−4
1.83×10−3
All vs. only geometry
+0.011
0.536
0.577
All vs. all w/o social
−0.006
0.374
0.436
All vs. all w/o text
+0.006
0.374
0.436
Only geometry vs. only text
+0.628
3.66×10−4
1.17×10−3
Only geometry vs. only social
+0.551
1.39×10−4
5.74×10−4
Appendix
Table 13: Paired fold-level F1 comparisons for the selected ablation settings. Positive ΔF1 means that the first condition outperforms the second. Dashes indicate identical fold-level F1 values or degenerate zero differences, for which the paired t -test is not defined.
k
Retained
Positive
Negative
F1
Recall
1
2134
1056
1078
0.793
0.773
2
1390
605
785
0.871
0.865
3
965
387
578
0.922
0.931
4
656
240
416
0.923
0.910
5
465
163
302
0.942
0.939
6
321
107
214
0.981
0.981
Appendix
Table 14: XGBoost results on the extended corpus.
Corpus
Pos.
Neg.
W=5 d
W=7 d
W=14 d
W=21 d
Base
HydroNewsFr
42
211
0.776 [0.67, 0.86]
0.762 [0.65, 0.86]
0.780 [0.68, 0.87]
0.696 [0.60, 0.79]
0.285
ClimateNewsFr
26
316
0.821 [0.70, 0.92]
0.800 [0.66, 0.91]
0.737 [0.59, 0.85]
0.778 [0.64, 0.89]
0.141
HydroNewsFr Ext.
131
345
0.875 [0.83, 0.92]
0.885 [0.84, 0.92]
0.874 [0.83, 0.91]
0.873 [0.83, 0.91]
0.432
Appendix
Table 15: Forward-chaining results at TA , for test windows of W days. Ext. denotes the extended hydrogen corpus of Appendix D . Cells report pooled F1 with a 95% percentile bootstrap interval; Base is the constant-positive baseline. Counts are the pooled evaluated articles at W=7 d and vary slightly with W under the window-fit rule.
In studies of media coverage of extreme climate events, NLP methods have become indispensable for identifying relevant texts in large news databases. Still, enough annotated data to train accurate deep learning-based classifiers from scratch is often not available. Topic Models have the advantage of being both unsupervised and interpretable, but are typically used only for exploratory analysis or data characterisation. In this study, we investigate how to employ Topic Models as binary classifiers for refining the retrieval of relevant news about seven types of extreme climate events in the German media. Our method relies on the posterior distributions estimated by Topic Models to select relevant documents, without modifying their training procedure. Using an annotated sample to guide the evaluation, we show that the probabilities assigned to keywords used to query news databases can also be informative for selecting relevant topics and improve sample precision. We compare our results to a fine-tuned text embedding classifier and an open-weight LLM, discussing observed trade-offs, e.g. the LLM's lowest precision. Moreover, we show that results are hazard-dependent, which speaks against considering climate events as a single category in NLP tasks.
Brielen Madureira, Mariana Madruga de Brito, Andreas Niekler
LeipzigLab - Climate Discourse, Leipzig University, Germany · Helmholtz Centre for Environmental Research - UFZ, Germany · Leipzig University, Germany +1
A social highlighter's most useful signal -- which passages a crowd of readers marks -- exists only for documents people have already read. Can the aggregate crowd salience of a document be predicted from its text before its marks accumulate? Prior work on this data found that zero-shot language models recover highlight locations worse than a trivial lead (position) baseline, so we ask whether a model trained on the highlight corpus can beat that baseline. Using a pre-registered ladder of models and a by-document cluster bootstrap, we find a small but robust edge: a logistic ranker over sentence embeddings and positional/contextual features beats the lead baseline by +0.044 average precision (95% CI [+0.029, +0.058]; clears a pre-registered margin delta=0.03 in 97% of resamples, and stable across pipeline re-runs). Two unsupervised extractive baselines (centroid, LexRank-style centrality) lose to lead, and the trained model beats them by +0.108, so the edge is not recovered by generic unsupervised proxies -- it reflects learning from real reader marks. In product terms, precision@3 rises from 0.25 to 0.39 (+55% relative) and the model beats lead on 69% of documents. An ablation attributes the edge to the raw embedding (+0.014) and training augmentation (+0.010), each with a positive CI. The edge is not a temporal-generalization failure, and we find no evidence that content drift or near-duplicate leakage explains it. A standardized regression shows the advantage is governed mainly by document popularity (lower popularity, larger edge) and by label reliability. It nearly vanishes only on the most popular content; there it is the lead baseline that strengthens, not the model that weakens. Because our evaluation conditions on documents that eventually accumulated readers, these results are a retrospective cold-start simulation.
Traditional topic modeling assigns a single topic to each document. In practice, however, many real-world documents, such as product reviews or open-ended survey responses, contain multiple distinct topics. This mismatch often leads to topic contamination, where unrelated themes are merged into a single topic, making it difficult to identify documents that truly focus on a specific subject. We address this issue by introducing segment-based topic allocation (SBTA), a reformulation of topic modeling that assigns topics not to entire documents, but to segments: short, coherent spans of text that each express a single theme. By modeling topical structure at the segment level, our approach yields cleaner and more interpretable topics and better supports analysis of multi-theme documents. To support systematic evaluation, we construct a SemEval-STM, a new dataset inspired by aspect-based sentiment analysis. Documents are first decomposed into topical segments using large language models (LLMs), followed by human refinement to ensure segment quality. We also propose a segment-level extension of the word intrusion task, enabling human evaluation of topical coherence at the granularity where topics are actually assigned. Across multiple models and evaluation metrics, we show that SBTA improves clustering quality and interpretability. Overall, this work provides a practical, scalable framework for fine-grained topic analysis in heterogeneous text corpora where documents naturally span multiple topics. URL: https://huggingface.co/datasets/LG-AI-Research/SemEval-STM
Hoonsang Yoon, Takyoung Kim, Wonkee Lee +3
1LG AI Research · University of Illinois Urbana-Champaign