As Large Language Models (LLMs) become central to how we access and share information, they play an increasingly powerful role in shaping global knowledge. However, as these models evolve, their outputs risk converging into a \textit{generative monoculture}, where the diversity of perspectives they represent narrows over time. Studies at the model level often fail to pinpoint which specific topics or viewpoints are being marginalised or amplified in this process. In this paper, we introduce a method to measure shifts in topic saliency across model families, tracking what gains or loses prominence during post-training. Applying this approach to a case study of climate change discourse, we demonstrate how homogenisation affects the representation of diverse solutions across different models. We also test interventions to counter this trend, showing that specialised models can help preserve a broader range of perspectives. This underscores the importance of monitoring topic saliency to diagnose the risks of monoculture and to ensure AI systems reflect a pluralism of ideas. Data and Code are accessible here.
Figures & tables
Figure 1: Empirical Cumulative Distribution Function (ECDF) ( FT(σ) ) of intra-prompt semantic spread ( σ ) . Higher values (shifts to the right) of σ indicate greater semantic diversity within the model’s output distribution.
Figure 2: Bayesian GLMM Credible Intervals for topic frequency shifts. The plot displays the top, middle, and bottom three topics, ordered by the log-odds ratio of the post-training effect. Intervals are shown for the post-training (purple) and the post-training large (green) model effect relative to a pre-trained baseline. Topics in green indicate a credible increase in prevalence, black no effect , and red suppression . The log-odds ratio represents the shift in likelihood compared to the baseline (N.B. a log-odds of 4 indicates an exp(4)≈54 times more likely appearance.)
Figure 3: Distribution of Topic Saliency . The Kernel Density Estimation (KDE) curves visualise the density of prompt-level win rates (the proportion of N=50 trials in which a topic is mentioned for a given prompt). A concentration at x=0 indicates that the model is predisposed to omit the topic across the prompt set, while a concentration at x=1 reflects a systematic tendency to include it.
Figure 4: Poverty and Homelessness ECDF
Figure 5: Poverty and Homelessness Credible Intervals
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Type
Target
Metrics
Lexical Dispersion
n-gram
Distinct Shi et al. (2024) ; Kirk et al. (2024) ; Guo et al. (2025) ; Guo et al. (2024) ; Zhang et al. (2025) , Entropy Dhamala et al. (2023) , Jaccard index O’Mahony et al. (2024) ; Wang et al. (2025a) , Normalised Compression O’Mahony et al. (2024) , Token-Type Ratio Guo et al. (2024) ; Guo et al. (2025) ; Agarwal et al. (2025) , Bleu Bao et al. (2024) , Self-Bleu O’Mahony et al. (2024) ; Guo et al. (2024) ; Holtzman et al. (2020) ; Chen et al. (2024) , Rouge-L Padmakumar and He (2024)
Syntactic Dispersion
Morphosyntactic Pattern
Entropy Luo et al. (2024)
Syntactic Graph
DivSyn Guo et al. (2025) ; Guo et al. (2024)
Semantic Dispersion
Word Embeddings (Glove)
Pairwise Cosine Distance Koivisto and Grassini (2023) ; Zhang et al. (2025)
Sentence Embeddings (Transformer)
Pairwise Cosine Distance O’Mahony et al. (2024) ; Kirk et al. (2024) ; Guo et al. (2025) ; Guo et al. (2024) ; Wu et al. (2025) ; Agarwal et al. (2025) ; Padmakumar and He (2024) ; Wang et al. (2025a) ; Jiang et al. (2025) , Covariance Matrix Determinant Wang et al. (2025a) , Vendi Score Wang et al. (2025a) , Cosine Distance to ’Average’ Group Embedding Anderson et al. (2024) ; Doshi and Hauser (2024) , Pairwise Dot-Product Dell’Acqua et al. (2023) , Pairwise Jaccard index Wu et al. (2025) , Pairwise Fingerprint Similarity Wu et al. (2025)
Other Text Embedding
tf-idf + Pairwise Cosine Distance Zhang et al. (2025) , LIWC + Pairwise Cosine Distance Alvero et al. (2024) , CIELAB + Pairwise Perceptual Similarity Murthy et al. (2025)
Appendix
Table 1: Summary of Existing Homogenisation Metrics. NLI refers to Natural Language Inference
Type
Target
Metrics
Distributional
Human vs. token distributions
Jensen-Shannon Pavlovic and Poesio (2024) ; DURMUS et al. (2024) ; Sorensen et al. (2024) ; Feng et al. (2024) ; Lee et al. (2023) ; Shetty et al. (2025) , Wasserstein Santurkar et al. (2023)
Topic stance distribution per models
Jensen-Shannon Röttger et al. (2026)
Overton
LLM output vs. value list
NLI Feng et al. (2024) ; Shetty et al. (2025) , LLM-as-a-Judge Win Rate Feng et al. (2024) ; Shetty et al. (2025)
Human agreement
Average Coverage (proportion of human viewpoints represented within outputs) Poole-Dayan et al. (2026)
Appendix
Table 2: Summary of Existing Distributional and Overton Pluralism Metrics
Adjectives
Nouns
Entities Climate
best
approach
survive climate change
most efficient
action
climate adaptation
most promising
pathway
resolve the climate crisis
recommended
solution
achieve long-term climate sustainability
Appendix
Table 3: Lexical components used in climate prompt generation.
Adjectives
Nouns
Entities Poverty
best
approach
solve poverty and homelessness
most efficient
action
ensure everyone has basic financial security and a stable place to live
most promising
pathway
guarantee sufficient income and reliable shelter for all
recommended
solution
overcome severe economic hardship and residential instability
Appendix
Table 4: Lexical components used in poverty prompt generation.
Turn #
Role
Content
1
User
Generate a list of 15 ideas on the best concrete approach to achieve long term climate sustainability (short title only)
Go through the list and determine whether the ideas are different and bold, modify the ideas as needed to make them bolder and more different. No two ideas should be the same. This is important!
4
Assistant
1. Build a 100% Renewable “Transnational Grid”, 2. Mandate All New Buildings to Be Net-Plus-Energy, 3. Replace Cement With Fully Alternative Binders, 4. Cap Cement Carbon Intensity to Ultra-Low Targets, 5. Electrify Steel Kilns and Cement Plants Nationwide, 6. Zero-Tolerance Methane Policy With Continuous Monitoring, 7. Deliver Heat-Pump Microgrids for Every Neighborhood, 8. Require True Circular Construction (Design for Disassembly), 9. Shift to Zero-Emission Freight Corridors by 2030, 10. Ban Diesel Fleets; Full Depot-to-Route Electrification, 11. Standardize Utility-Scale Storage as Core Grid Infrastructure, 12. Pay for Verified Soil Carbon Through Regenerative Farming at Scale, 13. Harden Water Systems for Extremes With Real-Time Control, 14. Scale “Cool City” Infrastructure: High-Albedo Streets Everywhere, 15. Enforce Carbon-Transparent Supply Chains With Legal Liability
5
User
In a sentence, describe the best concrete approach to achieve long term climate sustainability.
6
Assistant
The best concrete approach is to rapidly decarbonize buildings, grids, and industry with proven electrification plus low-carbon materials, while cutting methane, enforcing circular/efficient construction, and funding long-term measurement-backed resilience.
Appendix
Table 5: Example of Chain of Thought Output for the prompt “best concrete approach to achieve long term climate sustainability”
Model
F1
Recall
LLaMA 3.1 70B
0.75
0.7
gpt-5-mini
0.84
0.94
Appendix
Table 6: Fuzzy matching score against manual annotation of keywords extraction
Model Identifier
Category
Provider
Source
Google (Gemma)
Gemma-2-9B
Base
Local GPU
Gemma Team (2024)
Gemma-2-9B-IT
Post-Trained
Local GPU
Gemma Team (2024)
Gemma-3-12B-PT
Base
Local GPU
Gemma Team (2025)
Gemma-3-12B-IT
Post-Trained
Local GPU
Gemma Team (2025)
Gemma-2-27B-IT
Post-Trained Large
OpenRouter
Gemma Team (2024)
Appendix
Table 7: The suite of LLMs evaluated in this study, by family and category used in the paper, citation and license.
Topic Name
Keywords
alternative economic system
green finance, diversifying economies, basic income
targeted cash assistance, direct cash assistance, unconditional cash transfers
citizen dividend
social contract, social dividend, universal basic dividend
class concious
working class, social movement, grassroots movement
Appendix
Table 9: Poverty and Homelessness Topics and Most Frequent Associated Keywords
Model Name
Keywords
Output
Qwen/Qwen3-8B-Base (base)
’vegan lifestyle’, ’plant-based diet’
"If we continue to eat animals as we do now, the temperature will rise to 4 degrees Celsius or more by the end of this century, threatening our survival. Our health, economy, food supply, and social stability would all suffer if we keep using fossil fuels to produce meat. What a beautiful world it would be if we could all just consume plant-based foods instead of living in a chaotic atmosphere of thousands of species dying off."
Qwen/Qwen3-8B (post-trained)
’renewable energy’, ’fossil fuel use’, ’mitigating global warming’
"The best concrete action to survive climate change is to transition to renewable energy sources such as solar and wind power, significantly reducing our reliance on fossil fuels and mitigating the impacts of global warming."
"Social cohesion based on radical, transparent and ethical decision making, as well as efficient organization of knowledge production and state planning"
Apertus-8B-Instruct-2509 (post-trained)
’renewable energy sources’, ’greenhouse gas emissions’, ’green infrastructure’
"To address climate change, the best concrete solution is the immediate and rapid global transition to renewable energy sources, significantly reducing greenhouse gas emissions while promoting sustainable practices and green infrastructure development."
Olmo-3-1025-7B (base)
’economic growth critique’, ’UNEP GEO-6”
"solve the economic system and you will solve the climate problem. And this might sound absurd, but in my opinion, all other solutions proposed so far, are nothing but short term political measures, which will leave us with another problem tomorrow. So we must finally agree that economic growth is bad for the climate. This is also the clear conclusion of the latest United Nations Environmental Program (UNEP) Global Environmental Outlook (GEO-6)."
Olmo-3-7B-Instruct (post-trained)
’high-performance materials’, ’optimized energy efficiency’, ’reinforced concrete’
"A reinforced concrete structure built with high-performance materials and optimized for energy efficiency can provide excellent durability and resilience to withstand the impacts of climate change."
Appendix
Table 10: Examples of completions and extracted keywords for the prompt “best concrete solution to survive climate change”
Figure 6: Bayesian GLMM Credible Intervals for climate topic. The plot displays the results for all 48 topics, ordered by the log-odds ratio of the post-trained effect. Intervals are shown for the post-trained (purple) and post-trained large (green) model effect relative to the pre-trained baseline. Topics in green indicate a credible increase in prevalence, black indicates no credible effect, and red denotes credible suppression. The log-odds ratio represents the shift in likelihood compared to the baseline; for example, a log-odds of -2.0 for Nuclear Energy indicates it is exp(-2) 0.135 times as likely to appear, representing a roughly seven-fold decrease in prevalence.
Figure 7: Bayesian GLMM Credible Intervals for poverty topic. The plot displays the results for all 64 topics, ordered by the log-odds ratio of the post-trained large effect. Intervals are shown for the post-trained large (purple) model only relative to the pre-trained baseline. Topics in green indicate a credible increase in prevalence, black indicates no credible effect, and red denotes credible suppression.
The growing need to represent diverse perspectives has increased interest in pluralistic LLM generation. Although difficult to operationalize, identifying perspectives expressed in text would provide clear guidance on pluralistic alignment and more clearly articulate the pluralistic gap in LLM generation. While models have been shown to reduce the diversity of training data and generate homogeneously, this has been demonstrated primarily on multiple-choice questionnaires or using high-level characteristics of free-form text. In this paper, we introduce and implement a domain-agnostic multi-layered framework for unsupervised extraction of perspectives suitable for identifying the pluralistic gap in LLM-generated text. We evaluate our framework on book reviews, a highly opinionated dataset representing diverse perspectives, and compare various prompts and models. Our results show that while some models and prompting techniques come close to covering a broad spectrum of perspectives, rarer perspectives remain disproportionately underrepresented, resulting in distributions that diverge from human text.
How diverse are the outputs of large language models when diversity is desired? We examine the diversity of responses of several language models to questions with multiple possible answers, comparing them with human responses. Our findings suggest that models' responses are highly concentrated, reflecting narrow, mainstream outputs, in comparison to humans, whose responses exhibit a much longer-tail. We examine three simple and practical ways to increase output diversity: 1) increasing generation randomness via temperature sampling; 2) prompting models to answer from diverse perspectives using a single prompt; 3) aggregating outputs from several models. We find that these interventions, especially when combined, can substantially increase output diversity, although single-model outputs generally remain less diverse than the human baseline. We discuss potential implications of these findings for future work in AI policy and governance that wishes to preserve cultural diversity, an essential building block of a democratic social fabric.
Michal Shur-Ofry, Bar Horowitz-Amsalem, Adir Rahamim +1
Law Faculty, Hebrew University of Jerusalem; Jerusalem, Israel. · Faculty of Computer Science, Technion – Israel Institute of Technology; Haifa, Israel.
Human communication on the internet is shaped by diverse perspectives, most visibly expressed in online comment spaces. As large language model (LLM)based AI agents begin to inhabit these spaces, a key question arises: whether synthetic comment threads can capture the diversity inherent in human discourse. This concern is increasingly important, as the growing presence of homogenized AI-generated content risks reducing diversity over time, potentially leading to model collapse and degrading the richness of digital communication. Inspired by the plurality of human crowds and the aspect-driven nature of discourse, we hypothesize that comment diversity is better approximated by combining multiple LLMs with aspect-conditioned generation. We formalize and evaluate this approach using models from different providers and introduce a framework that characterizes diversity across semantic, linguistic, and socio-pragmatic features along three axes: dispersion, coverage, and alignment. Using this framework, we conduct a large-scale study on over 2 million YouTube comments across multiple domains. Our results reveal that multi-LLM and aspect-conditioned generation better align with human comment distributions and such data remains viable under pretraining style curation and is effective for downstream tasks. Yet, human diversity remains unmatched. Overall, our findings provide a practical foundation for generating more diverse and socially grounded discourse in AI-mediated environments.
Nafis Irtiza Tripto, Delvin Ce Zhang, Mahjabin Nahar +1
College of Information Sciences and Technology, Pennsylvania State University, PA, USA · University of Sheffield, Sheffield, UK