One Model Is Not a Crowd: Multi-LLM and Aspect-Conditioned Diverse Comment Generation
Organizations: College of Information Sciences and Technology, Pennsylvania State University, PA, USA · University of Sheffield, Sheffield, UK
Abstract
Human communication on the internet is shaped by diverse perspectives, most visibly expressed in online comment spaces. As large language model (LLM)based AI agents begin to inhabit these spaces, a key question arises: whether synthetic comment threads can capture the diversity inherent in human discourse. This concern is increasingly important, as the growing presence of homogenized AI-generated content risks reducing diversity over time, potentially leading to model collapse and degrading the richness of digital communication. Inspired by the plurality of human crowds and the aspect-driven nature of discourse, we hypothesize that comment diversity is better approximated by combining multiple LLMs with aspect-conditioned generation. We formalize and evaluate this approach using models from different providers and introduce a framework that characterizes diversity across semantic, linguistic, and socio-pragmatic features along three axes: dispersion, coverage, and alignment. Using this framework, we conduct a large-scale study on over 2 million YouTube comments across multiple domains. Our results reveal that multi-LLM and aspect-conditioned generation better align with human comment distributions and such data remains viable under pretraining style curation and is effective for downstream tasks. Yet, human diversity remains unmatched. Overall, our findings provide a practical foundation for generating more diverse and socially grounded discourse in AI-mediated environments.
Figures & tables
| Domain | # Ch. | # Vid. | # Comm. | # Words | Asp./Vid. |
| News | 14 | 3,857 | 518.7K | 10.27M | |
| Pop | 22 | 2,203 | 326.5K | 4.56M | |
| Tech | 10 | 1,514 | 269.0K | 5.34M |
| Axis | Definition | Semantic | Lexical | Pragmatic |
| Dispersion | Intra-set variation within comments | Clipped pairwise cosine dissimilarity (*, ) ( Cann et al., 2023 ; Guo et al., 2025 ; Zhang et al., 2025a ) | N-gram diversity ( ) ( Padmakumar and He, 2024 ) Self-repetition score ( ) ( Salkar et al., 2022 ) POS Compression ratio ( ) ( Shaib et al., 2025 ) Homogenization Score ( ) ( Lin, 2004 ) | Simpson diversity index (*, ) ( Simpson, 1949 ) |
| Coverage | Fraction of human patterns recovered | Weighted manifold recall (*, ) ( Kynkäänniemi et al., 2019 ) | Weighted n-gram & POS coverage (*, ) | Categorical coverage ratio (*, ) |
| Alignment | Distribution similarity between and | Maximum Mean Discrepancy (RBF kernel) ( ) | Avg. JSD score from ( ) | Avg. category JSD ( ) |
Appendix figures & tables36 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Description |
|---|---|
| Content and Comment Space | |
| A content item (video). | |
| The space of all possible comments. | |
| A comment, modelled as a random variable. | |
| Fixed embedding function mapping comments into a -dimensional semantic space. | |
| A set of human-generated comments for video . | |
| Domain | #Channels | #Videos | #Comments | Channels (clickable, ordered by #videos) |
| News | 14 | 3,857 | 518,718 | Fox News (663, 134,427); ABC News (572, 78,937); CBS News (457, 47,384); MSNBC (379, 55,935); BBC News (379, 42,824); NBC News (275, 25,079); DW News (248, 19,516); CNBC (215, 29,415); Sky News (178, 15,488); CNN (169, 34,187); Al Jazeera (157, 15,315); Wall Street Journal (143, 18,314); AFP News (15, 1,124); Associated Press (7, 773) |
| Pop | 22 | 2,203 | 326,542 | Nintendo (469, 78,785); PlayStation (405, 63,392); Netflix (363, 58,753); GameSpot Trailers (253, 32,632); Xbox (189, 17,743); Sony Pictures (60, 8,754); Netflix India (53, 7,878); Prime Video (51, 7,470); Marvel (48, 8,164); Warner Bros. (46, 5,779); Universal (41, 5,910); GameTrailers (39, 4,967); 20th Century (31, 4,343); Hulu (31, 3,479); Paramount (28, 4,171); DC (24, 4,204); IGN (22, 3,915); Prime India (20, 3,201); Epic Games (13, 1,168); Prime UK (10, 1,129); Ubisoft (5, 540); Apple TV (2, 165) |
| Tech | 10 | 1,514 | 268,954 | Linus Tech Tips (272, 45,354); Unbox Therapy (207, 30,910); CNET (201, 25,949); Mrwhosetheboss (158, 36,402); Android Authority (135, 19,224); Austin Evans (131, 31,866); GSMArena (112, 17,750); Dave2D (109, 22,779); UrAvgConsumer (103, 28,077); MKBHD (86, 10,643) |
| Aspect | % | Pragmatic | Aspect Description |
| Enthusiastic Fan Reactions (Pragmatic) | 65.07 | informal joy non-sarcastic | Strong excitement around the film and its Bruce Springsteen connection, with nostalgic references to the soundtrack and repeated praise for the trailer, alongside expressive and hyperbolic fan reactions. |
| Music-driven Emotional Resonance (Topic) | 13.97 | – | Viewers share emotional connections to Springsteen’s music and highlight how the film’s coming-of-age narrative resonates with personal experiences and nostalgia. |
| Emotional & Relatable Storytelling (Topic) | 13.10 | – | Comments emphasize the film’s emotional depth, relatability, and themes of family, identity, and struggle, often describing cathartic and uplifting viewing experiences. |
| Springsteen Legacy Appreciation (Topic) | 4.80 | – | Discussion centers on admiration for Springsteen’s influence, with personal reflections on how his music shapes identity and enhances the film’s impact. |
| Cultural Identity & Representation (Topic) | 2.18 | – | Comments focus on South Asian and working-class identity, highlighting cultural pressures and the empowering role of music in shaping self-expression. |
| Playful Music Culture Observations (Pragmatic) | 0.87 | sentiment mixed light humor | Lighthearted remarks about music preferences and cultural trends, combining casual humor with appreciation for classic rock influences. |
| Provider | Models |
| OpenAI (7) | ✓ gpt-oss-20b |
| ✓ gpt-oss-120b | |
| ✗ gpt-5 | |
| ✗ gpt-5-mini | |
| ✗ gpt-5-nano | |
| ✗ gpt-4.1 |
| Dataset | H | S | M | Comment | |||
| News | H M S | ||||||
| Pop | H M S | ||||||
| Tech | H M S | ||||||
| All | H M S |
| Dataset | Metric | S | M | Comment | |||
| News | Manifold recall | ||||||
| News | Semantic recall (MaxSim) | ||||||
| News | Semantic recall @ | ||||||
| Pop_culture | Manifold recall | ||||||
| Pop_culture | Semantic recall (MaxSim) | ||||||
| Pop_culture | Semantic recall @ |
| Dataset | S | M | Comment | |||
| News | ||||||
| Pop_culture | ||||||
| Tech | ||||||
| All |
| Dataset | Metric | H | S | M | Comment | |||
| News | n-gram diversity | H M S | ||||||
| News | TTR | H M S | ||||||
| News | Self-repetition | H M S | ||||||
| News | Compression ratio | H M S | ||||||
| News | POS comp. ratio | H M S | ||||||
| Pop_culture | n-gram diversity | H M S |
| Dataset | S | M | Comment | |||
| News | M S | |||||
| Pop_culture | M S | |||||
| Tech | M S | |||||
| All | M S |
| Dataset | S | M | Comment | |||
| News | M S | |||||
| Pop_culture | M S | |||||
| Tech | M S | |||||
| All | M S |
| Dataset | H | S | M | Comment | |||
| News | M H S | ||||||
| Pop_culture | M H S | ||||||
| Tech | M H S | ||||||
| All | M H S |
| Dataset | S | M | Comment | |||
| News | ||||||
| Pop_culture | ||||||
| Tech | ||||||
| All |
| Dataset | S | M | Comment | |||
| News | ||||||
| Pop_culture | ||||||
| Tech | ||||||
| All |
| Feature | Axis/Metric | Dataset | H | M | ||
| Semantic | Dispersion | News-new | ||||
| Pop-culture-new | ||||||
| Tech-new | ||||||
| Semantic | Manifold recall | News-new | – | |||
| Pop-culture-new | – | |||||
| Tech-new | – |
| Metric | Krippendorff’s | Weighted Cohen’s | Agreement Level |
| Naturalness | 0.601 | 0.591 | Moderate |
| Authenticity | 0.463 | 0.471 | Moderate |
| Overall Quality | 0.447 | 0.446 | Moderate |
| Representativeness | 0.413 | 0.360 | Fair |
| Diversity | 0.369 | 0.397 | Fair |
| Appropriateness | 0.298 | 0.312 | Fair |
| Metric | S | M | H | Comment | |
| Diversity | H M S | ||||
| Context | H M S | ||||
| Naturalness | H M S | ||||
| Authenticity | H M S | ||||
| Representativeness | H M S | ||||
| Appropriateness | H M S |
| Metric | Human Judgment | LLM Judgment | Setting-level | Sample-level |
| Diversity | H M S | H M S | 1.00 | 0.16 |
| Context | H M S | S M H | -0.63 | 0.04 |
| Naturalness | H M S | H M S | 1.00 | 0.07 |
| Authenticity | H M S | H M S | 1.00 | 0.21 |
| Representativeness | H M S | H M S | 1.00 | 0.43 |
| Appropriateness | M S H | S M H | 0.63 | 0.05 |
| Setting | Anger | Disgust | Fear | Joy | Neutral | Sadness | Surprise | Mac.F1 | ||||||||||||||
| P | R | F1 | P | R | F1 | P | R | F1 | P | R | F1 | P | R | F1 | P | R | F1 | P | R | F1 | ||
| H | 0.55 | 0.56 | 0.55 | 0.57 | 0.53 | 0.55 | 0.73 | 0.34 | 0.46 | 0.77 | 0.80 | 0.79 | 0.84 | 0.86 | 0.85 | 0.57 | 0.56 | 0.56 | 0.72 | 0.71 | 0.71 | 0.64 |
| S | 0.44 | 0.56 | 0.49 | 0.45 | 0.60 | 0.52 | 0.68 | 0.46 | 0.55 | 0.79 | 0.73 | 0.76 | 0.82 | 0.83 | 0.82 | 0.62 | 0.57 | 0.59 | 0.68 | 0.64 | 0.66 | 0.63 |
| M | 0.42 | 0.64 | 0.51 | 0.51 | 0.43 | 0.47 | 0.73 | 0.39 | 0.51 | 0.78 | 0.74 | 0.76 | 0.82 | 0.84 | 0.83 | 0.61 | 0.54 | 0.58 | 0.71 | 0.67 | 0.69 | 0.62 |
| 0.53 | 0.55 | 0.54 | 0.54 | 0.52 | 0.53 | 0.75 | 0.41 | 0.54 | 0.74 | 0.81 | 0.77 | 0.84 | 0.83 | 0.84 | 0.63 | 0.53 | 0.57 | 0.68 | 0.71 | 0.70 | 0.64 | |
| 0.57 | 0.56 | 0.57 | 0.56 | 0.50 | 0.53 | 0.81 | 0.43 | 0.56 | 0.75 | 0.77 | 0.76 | 0.82 | 0.85 | 0.83 | 0.62 | 0.49 | 0.54 | 0.68 | 0.71 | 0.69 | 0.64 | |
| Conservative | Liberal | Neutral | Overall | |||||||||
| Model | Condition | P | R | F1 | P | R | F1 | P | R | F1 | Acc | Mac.F1 |
| Qwen-0.5B | Few-shot (H) | 0.48 | 0.36 | 0.41 | 0.41 | 0.59 | 0.49 | 0.20 | 0.06 | 0.09 | 0.43 | 0.33 |
| Few-shot ( ) | 0.51 | 0.19 | 0.28 | 0.42 | 0.72 | 0.53 | 0.22 | 0.23 | 0.22 | 0.42 | 0.34 | |
| Mistral-7B | Few-shot (H) | 0.89 | 0.83 | 0.86 | 0.90 | 0.51 | 0.65 | 0.22 | 0.82 | 0.35 | 0.69 | 0.62 |
| Few-shot ( ) | 0.91 | 0.84 | 0.87 | 0.93 | 0.47 | 0.63 | 0.20 | 0.82 | 0.32 | 0.68 | 0.61 | |
| Gemma-12B | Few-shot (H) | 0.74 | 0.95 | 0.83 | 0.96 | 0.54 | 0.69 | 0.41 | 0.66 | 0.51 | 0.75 | 0.68 |