Iterative Topic Taxonomy Induction with LLMs: A Case Study of Electoral Advertising
Authors: Alexander Brady, Tunazzina Islam
Organizations: Department of Computer Science, ETH Zürich, Zurich, Switzerland · Department of Computer Science, Purdue University, West Lafayette, IN 47907, USA
Social media platforms play a pivotal role in shaping political discourse, but the scale and rapid evolution of online content make systematic analysis difficult. We introduce an end-to-end framework for inducing an interpretable topic taxonomy from unlabeled text corpora. The framework combines embedding-based clustering with iterative large language model (LLM) inference to construct a topic taxonomy without requiring predefined labels or seed topics. It first synthesizes candidate topics from document clusters and then uses the resulting taxonomy to assign consistent topic labels across clusters. We evaluate the approach through a case study of political advertising ahead of the 2024 U.S. presidential election. We use the induced taxonomy to support downstream analyses of issue prevalence, moral framing, advertising spend, and demographic exposure patterns. These results suggest that iterative taxonomy construction can provide a scalable and interpretable approach to organizing large unlabeled text corpora while supporting substantive downstream analysis.
Figures & tables
Figure 1: Overview of the proposed framework. Documents are first embedded and clustered. (a) Cluster representatives are processed sequentially by the LLM to induce a shared topic taxonomy. (b) The LLM assigns one topic label to each cluster from the completed taxonomy.
Issue
#Clusters
Example Ad
economy
15
Molly Buck’s agenda is failing Iowa families. We can’t afford Molly Buck in the State House.
voting rights
13
Make Your Vote Count. RE-register to Vote in the District of Your Second Home.
crime/justice
12
Orange County Firefighters Trust Dave Min To Keep Our Communities Safe.
personal freedom
4
Tim Sheehy will always fight for Montana Values in the Senate!
voting
4
Vote Rebecca for State Rep by Nov 5th!
education
4
Vote for students. Vote for teachers. Support our public schools.
Table 1: Topics identified by the LLM in the political ads dataset. ‘#Clusters’ column gives the number of unique clusters assigned to each topic by the LLM.
Model
Average Score
Best Label
BERTopic
1.2
3
TopicGPT-style
2.7
5
Our Method
2.8
12
Table 2: Annotation results. The average score is out of 5, and the best label is the number of times that model’s label was selected as the most fitting (among all annotators).
Model
Macro F1
Accuracy
Logistic Regression
0.37
0.55
XGBoost
0.31
0.53
RoBERTa
0.32
0.53
SetFit
0.36
0.60
Table 3: Document-level classification performance for topic assignment.
Figure 2: Associations between moral-foundation labels and topic labels for (a) the proposed taxonomy and (b) LDA topics. Cells show Pearson correlations between one-hot topic and moral-foundation assignments.
Figure 3: Average advertising spend and impressions per ad by annotated topic.
Figure 4: Spend by the top five funders of economy, crime/justice, and abortion ads, segmented by moral foundation.
Figure 5: Distribution of advertising spending across moral foundations within economy, crime/justice, and abortion advertisements. Percentages indicate within-topic spending shares; total spending is shown beneath each chart.
Figure 6: PPMI associations between political issue or moral-foundation labels and observed audience demographics. Higher values indicate that a label occurs disproportionately among advertisements delivered to the corresponding demographic group.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Care/Harm: It suggests that someone other than the speaker is worthy of compassion or is experiencing harm, grounded in the values of kindness, tenderness, and care.
Fairness/Cheating: Emphasizes justice, personal rights, and independence; involves comparing with other groups. Advocates for equal opportunity and resists those who benefit without contributing (“Free Riders").
Loyalty/Betrayal: Based on the values of loyalty to one’s country and willingness to sacrifice for the group. Activated by a sense of unity and collective responsibility—“one for all, and all for one".
Authority/Subversion: Centers on showing respect (or resistance) toward established authority and following long-standing traditions. It includes maintaining social order and fulfilling the duties tied to hierarchical roles, such as obedience, respect, and role-based responsibilities.
Sanctity/Degradation: Beyond religion, this value highlights respect for human dignity and aversion to moral or physical corruption, promoting purity, self-control, and the belief that the body is sacred and vulnerable to defilement.
Liberty/Oppression: Captures the feelings of reactance and resentment people experience when their freedom is restricted, often leading to collective disdain for authoritarian figures and motivating unity and resistance against oppression.
Appendix
Table 4: Six basic moral foundations Iyer et al. (2012) ; Haidt and Graham (2007) ; Haidt and Joseph (2004) .
A large part of political discourse on social media is built and expressed at a level of claims: i.e. declarative, typically single-clause statements, which convey a particular interpretation of reality and can range from factual to evaluative. Moreover, rather than occurring randomly, claims coalesce, recur in patterns, and come to be associated with different world views. When paired with structural computational tools such as Social Network Analysis, claims can be a powerful unit of analysis to study political phenomena such as echo chambers or polarisation. In this paper, we present a pipeline that uses a large language model (LLM) to inductively extract and catalogue claims from large social media corpora, and apply it to two different Twitter datasets: one relating to the 2020 US presidential election and the other to the 2022 FIFA World Cup. We comprehensively evaluate the approach by measuring the pipeline's recall and precision against manually annotated samples, run ablation studies isolating the contribution of its various components, and perform a qualitative error analysis. We discuss the value of the approach in the context of Computational Social Science research, and illustrate its capabilities by presenting the claims catalogue obtained from each dataset.
We present a new computational framework for detecting and structuring manipulative political narratives. A task that became more important due to the shift of political discussions to social media. One of the primary challenges thereby is differentiating between manipulative political narratives and legitimate critiques. Some posts may also reframe actual events within a manipulative context. To achieve good clustering results, we filter manipulative posts beforehand using a detailed few-shot prompt that combines documented campaign narratives with legitimate criticisms to differentiate them. This prompt enables a reasoning model to assign labels, retaining only manipulative narrative posts for further processing. The remaining posts are subsequently embedded and dimensionality-reduced using UMAP, before HDBSCAN is applied to uncover narrative groups. A key advantage of this unsupervised approach is its independence from a predefined list of target categories, enabling it to uncover new narrative clusters. Finally, a reasoning model is employed to uncover the narrative behind each cluster. This approach, applied to over 1.2 million social media posts, effectively identified 41 distinct manipulative narrative clusters by integrating prompt-based filtering with unsupervised clustering.
We ask whether topic sentiment has a causal effect on perceived political ideology, and whether the answer depends on who assigns the ideology label. Using articles from AllSides, paired with shared sentiment annotations from Llama-3.3-70b-versatile, we compare ideology labels from expert human annotators, GPT-4o-mini (baseline and finetuned), and Llama-3.3-70B. We apply Double Machine Learning (DML) and mediation analysis across all four annotation paradigms. Zero-shot LLMs regularly inflate effect sizes relative to human annotations, while fine-tuning often attenuates them back toward the human scale. Our results have implications for the use of LLM annotations as silver labels and as proxies for human judgment in downstream causal analyses: they may be reliable for recovering the presence and direction of effects on the partisan topics, but not their magnitude, leading to over- or under-prediction of some ideology given particular topics.
Upasana Chatterjee
Department of Computer Science Columbia University