Why Does Misinformation Propagate Faster? An Algorithmic Perspective on X
Organizations: Arizona State University
Abstract
Misinformation is widely reported to propagate faster on engagement-based platforms, yet prior work largely focused on empirical analysis, without identifying a specific algorithmic mechanism that results in this phenomenon. Thanks to the open-sourcing of X's recommendation algorithms, we conduct what is, to our knowledge, the first component-level study of the recommendation algorithm deployed by a social media platform, which examines how each of its components affects misinformation propagation. Specifically, we identify the engagement fungibility mechanism in the algorithm, where the final recommendation score is constructed as a weighted sum of all predicted user activities. As a result, a tweet can be repeatedly recommended simply because it is predicted to draw many instant reactions (e.g., likes and retweets), even when it is not expected to draw thoughtful responses (e.g., replies and quotes). Since misinformation typically draws a larger share of its engagement from instant reactions, this mechanism enables it to receive more recommendation exposure and to propagate faster. To empirically validate this mechanism, we re-implement X's recommendation algorithm on the USC X 2024 election corpus, and build a calibrated simulation study to analyze the impact of different scoring rules. We find that re-tuning the metric weights has little or even a negative impact on reducing the credibility exposure gap, while those scoring rules that set a precondition of thoughtful engagement for amplification would be able to alleviate the gap significantly, across 46 robustness checks. Our diagnosis, therefore, yields a simple and deployable fix, a reflective-threshold gate that withholds amplification until a tweet is predicted to draw thoughtful engagement, which we find to reallocate exposure away from low-credibility content at no cost to mainstream exposure and with no loss of engagement.
Figures & tables
| Research stream | Level of analysis | Method | Inside algorithm |
| Misinformation research (§ 2.1 ) | User and interface | Lab and survey experiments, analytical models | No |
| Algorithmic amplification (§ 2.2 ) | Platform output | Audits, field experiments | No |
| Recommender and simulation (§ 2.3 ) | Whole algorithm | Field experiments, simulation | No |
| This paper | Algorithmic component | Calibrated simulation of the open-sourced algorithm | Yes |
| Action | Cognitive effort | Dual-process class |
| Like | one click, no reading | fast (System 1) |
| Retweet | one click, often no reading, in-group signaling | fast (System 1) |
| Reply | read compose text defend a position | slow (System 2) |
| Quote-tweet | read compose frame for own audience | slow (System 2) |
| Mean | Median | |||
| Coverage | Engagement per tweet | |||
| Tweets | Replies | |||
| Unique authors | Retweets | |||
| Unique conversations | Likes | |||
| Quote-tweets | ||||
| Composition (% of tweets) | Impressions |
| Component | Population-level modeling | Per-agent state (static) |
| User pool | 50,000 users sampled from observed USC authors, stratified by follower-count quintile paid verification. | Follower/friend/status/favorite/list counts and the paid-verification flag |
| Seed population | Sampled from observed tweets, stratified by credibility label for the hypothesis test, each seed treated as a fresh cascade root. | Posting time, text, embedded URLs, and engagement counts (training targets only) |
| Ranker aggregation layer | Parallel MaskNet ( Wang et al. 2021 ) with four engagement objectives, published weights for replies, retweets, and likes, and the scoring rule as the only component varied. | Per-objective probabilities, aggregate and relative scores, exposure allocation, and activity indicator |
| Scoring rule | Formula | Role |
| Additive (baseline) | Production-style baseline | |
| Slow-gates-fast multiplicative (F1) | Reflective-gated form | |
| Ratio-correction (F2) | Reflective-gated form | |
| Reflective-threshold sigmoid (F3) | Reflective-gated form | |
| Retuned additive (control) | Additive with halved | Parameter-only control |
| Metric | Observed mean | Simulated mean | Welch | |
| Root reply count | ||||
| Aggregate engagement | ||||
| Reflective engagement share | ||||
| Time to peak (hours) |
| Statistic | Observed impressions | Simulated exposure |
| Low-credibility share of the total | ||
| Credibility gap in shares (low high) | pts | pts |
| Mean ratio, low to high | ||
| Gini across tweets | ||
| Top 10% share of the total | ||
| Scoring rule | Exposure contrast | Cascade size contrast |
| Slow-gates-fast multiplicative (F1) | ∗∗∗ (4.34) | ∗∗∗ (0.94) |
| Ratio-correction (F2) | ∗∗∗ (0.28) | ∗∗ (0.38) |
| Reflective-threshold sigmoid (F3) | ∗∗∗ (5.05) | ∗∗∗ (2.48) |
| Retuned additive (control) | ∗∗∗ (0.11) | ∗∗ (0.26) |
| Fidelity to the published ordering | Configurations | Best gap reduction attained |
| Ordering closely preserved ( ) | 1,489 | |
| Ordering loosely preserved ( ) | 1,946 | |
| Unconstrained | 2,083 |
| Feature | Reflective-to-reactive gap | Interpretation |
| Listed count (log) | prestige reply-heavy | |
| Followers (log) | prestige reply-heavy | |
| Text length | substance reply-heavy | |
| Paid verification | verification reply-heavy | |
| Follower-to-friend ratio (log) | influence reply-heavy | |
| Mention count | conversation reply-heavy |
| Metric (mean per tweet) | low-credibility ( ) | high-credibility ( ) |
| Replies | ||
| Retweets | ||
| Likes | ||
| Reply share of engagement | ||
| Cascade-size proxy |
| Quantity (mean per tweet) | low-credibility | high-credibility |
| Predicted reactive/reflective | ||
| Check | Variation | Finding |
| 1 | Aggregation form (F1–F3 and the retuned control) | All three reflective-gated forms narrow the gap significantly, the control moves it the opposite way |
| 2 | Slow/fast partition (four assignments) | Significantly negative under all four |
| 3 | Ranker training seed (five retrainings) | Contrasts cluster in , every one significant |
| 4 | Corpus period (adjacent week-long slice) | on the adjacent slice against at baseline |
| 5 | Partition training seed ( combinations) | Significantly negative in all 20 combinations |
| 6 | User-pool size (25K, 50K, 100K users) | Identical contrasts at all three sizes |
| Hypothesis | Sections | Result |
| H1: The aggregation form narrows the gap | 6.2 , 6.5 | Supported at both layers, on exposure under all three reflective-gated forms and in all 46 robustness variations, and on cascade size under the baseline ranker. |
| H2: Re-tuning the weights does not help | 6.2 , 6.3 | Supported, as the retuned control moves the gap in the opposite direction and no weight configuration reaches the threshold form’s gap reduction. |
| H3: The gate acts on the engagement mix | 6.4 | Supported, as demotion tracks the predicted reactive-to-reflective ratio, and the placebo gate and the correlation analysis show that neither alternative explanation accounts for the targeting. |
| Low-credibility exposure share | High-credibility exposure share | Expected engagement per exposure | Top-20% overlap | |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Training corpus | Metric | Scoring rule | Estimate | |
| Baseline | 1.0M | Exposure | F1 (multiplicative) | ∗∗∗ (4.34) |
| Baseline | 1.0M | Cascade size | F1 (multiplicative) | ∗∗∗ (0.94) |
| Enlarged | 2.0M | Exposure | F1 (multiplicative) | ∗∗∗ (3.93) |
| Enlarged | 2.0M | Cascade size | F1 (multiplicative) | ∗∗∗ (0.82) |
| Enlarged | 5.0M | Exposure | F1 (multiplicative) | ∗∗∗ (4.25) |
| Enlarged | 5.0M | Cascade size | F1 (multiplicative) | ∗∗∗ (0.73) |
| Subset | F1 score contrast | Control score contrast | |
| Full | 18 | ∗∗∗ | ∗∗∗ |
| Tweet and time only | 8 | ∗∗∗ | ∗∗∗ |
| Account only | 10 | ∗∗ | ∗∗∗ |
| Minimal | 3 | ∗∗∗ | ∗∗∗ |
| Form | Setting | Estimate |
| F1 (multiplicative) | ∗∗∗ (1.54) | |
| ∗∗∗ (2.70) | ||
| ∗∗∗ (4.34) | ||
| ∗∗∗ (8.24) | ||
| ∗∗∗ (13.0) | ||
| ∗∗∗ (18.4) |
| Iteration | Event model | Estimate of the dispersion | KS (root reply) | KS (reflective share) |
| (i) | Poisson with the mean matched | — | 0.60 | 0.62 |
| (ii) | Zero-inflated Poisson | — | 0.090 | 0.097 |
| (iii) | Zero-inflated Negative-Binomial | Mean and variance of the non-zero counts | 0.165 | 0.150 |
| (iv) | Zero-inflated Negative-Binomial | Mean and variance after dropping the largest 5% of the non-zero counts | 0.057 | 0.061 |