Limited review teams must identify which emerging claims are likely to keep growing before their eventual reach is known. We center early misinformation triage on this continuation-forecasting problem: predicting subsequent recorded propagation-tree growth from the first 30 minutes of activity. On FibVID, we compare early node count with structural depth entropy, temporal arrival entropy, and their pair while keeping all propagation trees from each original claim in one partition. Across 352 test trees from 59 claim groups separate from training, the combined model raises R2 for log-transformed future growth from 0.307 to 0.323 and reduces log-MAE by 2.4% (95% claim-bootstrap CI, -0.4% to 5.2%). The gain is especially pronounced among 97 high-activity trees: R2 rises from 0.248 to 0.395, Spearman's ρ from 0.394 to 0.529, and log-MAE falls by 11.7% (95% CI, -1.9% to 23.9%). Complementing the 30-minute growth forecast, we analyze the first 15 replies in 563 PHEME threads. In this cohort, the 15th reply arrives after a median of 28.8 minutes; 52.0% reach the fixed reply prefix within 30 minutes and 71.6% within one hour. Even without the LLM-generated factual-accuracy dimension, the remaining stance, communicative, and affective state composition retains cross-event ranking signal (ROC-AUC 0.538); including that dimension increases ROC-AUC to 0.562. In a separate Check-COVID evaluation of 229 claims, reciprocal-rank fusion retrieves a gold evidence document within the top five for 74.2% of claims and within the top 20 for 94.3%; sentence reranking reaches Recall@20 of 58.1%. We propose an integrated human-review system that brings these early forecasts, response patterns, and retrieved evidence together for misinformation triage relying on the potential virality of claims.
Figures & tables
Figure 1 : Proposed asynchronous early-triage workflow. The growth forecast is available at 30 minutes, while response context is added when 15 replies have arrived. The two predictive axes remain separate, and each contributes evidence to a human review queue.
Branch
Source
n
Unit
Evaluation role
Virality
FibVID
1,774
rooted propagation trees from 295 claims
1,422 training; 352 test trees, split by claim
Veracity
PHEME
563
binary-labeled threads with ≥15 raw replies
event-held-out screen; descriptive sparse LDA
Retrieval
Check-COVID
229
official test claims
gold-evidence retrieval evaluation
Table 1 : Datasets used in the reported experiments. “Eligible” indicates that all task-specific observation and label requirements were met.
Figure 2 : Illustration of structural depth and temporal bucketing in an early propagation tree.
Model
Log-MAE
R2
Spearman ρ
Node MAE
Count only
1.2323
0.3067
0.5401
75.60
Count + structural entropy
1.2139
0.3196
0.5514
74.10
Count + temporal entropy
1.2334
0.3039
0.5401
75.61
Count + both entropies
1.2029
0.3227
0.5535
72.21
Table 2 : FibVID propagation-tree continuation ablation on all 352 test trees from 59 claim groups. Log-MAE and R2 use log(1+Gc) ; node MAE uses the reconstructed number of subsequent recorded nodes. Errors are lower-is-better. Bold denotes the best value among the four configurations.
Model
Log-MAE
R2
Spearman ρ
Node MAE
Count only
0.6400
0.2482
0.3941
76.17
Count + structural entropy
0.5882
0.3738
0.5239
72.94
Count + temporal entropy
0.6253
0.2396
0.4327
74.31
Count + both entropies
0.5651
0.3954
0.5288
67.22
Table 3 : FibVID propagation-tree continuation ablation on 97 high-activity test trees from 35 claims, defined by nc30≥15 . Each model is fitted on the same complete training partition used in Table 2 . The subgroup is selected using only 30-minute activity.
Feature family
Candidates
ROC–AUC
PR–AUC
Bal. acc.
Macro-F1
Combined response, lexical, and structural
1,217
0.492
0.745
0.471
0.465
State composition
94
0.538
0.764
0.522
0.484
All response and lexical features
1,209
0.488
0.742
0.467
0.463
Prefix structure and timing
8
0.539
0.765
0.522
0.506
Entropy and transition dynamics
1,113
0.460
0.720
0.473
0.466
Table 4 : Primary PHEME event-grouped out-of-fold evaluation after excluding explicit factual-alignment-derived features. Each feature-family model is trained without the held-out events. Candidate counts are reported before within-fold variance filtering and sparse regularization. PR–AUC should be interpreted relative to the 73.5% true-class prevalence.
Figure 3 : Density of the 563 eligible PHEME threads projected onto the descriptive ten-feature sparse LDA axis. Positive values are oriented toward the true-associated distribution. Feature selection and fitting use this same complete sample, so the figure illustrates descriptive separation rather than held-out classification.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Event
False
True
Total
Sydney Siege
50
179
229
Ottawa Shooting
21
131
152
Charlie Hebdo
46
75
121
Germanwings crash
29
22
51
Ferguson
3
6
9
Gurlitt
0
1
1
Appendix
Table 6 : Event and veracity composition of the 563-thread PHEME sample.
Feature family
Candidates
ROC–AUC
PR–AUC
Bal. acc.
Macro-F1
State composition, without
94
0.538
0.764
0.522
0.484
State composition, with
104
0.562
0.774
0.539
0.501
All response and lexical, without
1,209
0.488
0.742
0.467
0.463
All response and lexical, with
1,305
0.543
0.765
0.545
0.532
Combined, without
1,217
0.492
0.745
0.471
0.465
Combined, with
1,313
0.546
0.767
0.549
0.536
Appendix
Table 7 : Event-held-out comparison without and with factual-alignment-derived features. Candidate counts are reported before within-fold variance filtering and sparse regularization.
Feature
Definition at the 30-minute observation cutoff
Observed node count
number of recorded nodes in the snapshot, including the root
Structural entropy
unnormalized Shannon entropy in bits over observed node depths, including the root; zero for a root-only snapshot
Temporal entropy
unnormalized Shannon entropy in bits over positive-time arrivals in six fixed five-minute bins; zero when no such arrival is observed
Appendix
Table 8 : The three predictors used in the FibVID propagation-tree continuation ablation. All are calculated at 30 minutes from the tree root.
Figure 4 : Predicted versus observed log(1+Gc) for the 97 high-activity FibVID test propagation trees. Left: count-only baseline. Right: count plus structural and temporal entropy. Both models are fitted on 1,422 trees from 236 other claims. The dashed line denotes perfect prediction.
Figure 5 : Check-COVID evidence-document retrieval (left) and evidence-sentence reranking (right). The dashed line is the maximum sentence recall permitted by the 20-document candidate set.
Purdue University West Lafayette, Indiana, USA · St John’s University New York, New York, USA · Stevens Institute of Technology Hoboken, New Jersey, USA +2