Open-ended LLM homogeneity can create false plurality when several systems appear to offer independent perspectives while returning the same familiar default. Single-pass answers obscure the distinction between agreement produced by a tightly constrained answer space, prompt-vocabulary echo, and broader answer spaces with stable alternatives beneath the surface. We introduce CHOIR (Collective Hierarchically-Ordered Inquiry Responses), a framework that adapts free-list elicitation from cognitive anthropology to LLM ensembles. CHOIR repeatedly elicits ranked lists, clusters items into prompt-level concepts, and measures concept salience across models, prompt variants, and persona conditions. We evaluate CHOIR on Infinity-Chat 100, an external prompt bank from recent work on open-ended model homogeneity, and on a 27-question targeted diagnostic bank designed to isolate mechanism-level contrasts. On Infinity-Chat 100, CHOIR reproduces high surface agreement (93/100 prompts above chance) while separating narrow prompts from broad prompts with recoverable depth. Across targeted probes and the external prompt bank, base-model identity remains the strongest recoverable signature, and persona prompts shift surfaced concepts within base-model signatures. A source-blind ranking module prioritises rare-but-stable candidates for later inspection. CHOIR turns open-ended homogeneity into a diagnostic measurement problem by asking where models converge, why they converge, and what remains reachable under structured depth probing.
Figures & tables
Figure 1: CHOIR experimental setup. Two prompt banks and model conditions feed a common elicitation-and-analysis core. Persona prompts alter the elicitation conditions. Source-blind ranking triages rare, stable candidates after measurement. The analysis distinguishes shared defaults, prompt echo, model signatures, and stable alternatives.
Empirical component
Purpose
Interpretation
Targeted probes
Provide controlled contrasts for vocabulary scaffolding, welfare/distress, conversational memory, computational self-description, and persona conditioning.
The questions are transparent study probes for identifying prompt echo, model-dominant signatures, and uneven persona-conditioned salience shifts.
Infinity-Chat 100
Test portability on an independently proposed prompt bank directly in the lineage of open-ended model homogeneity.
CHOIR generalises beyond targeted probes and frames homogeneity in terms of prompt width. Some prompts are narrow, whereas others have recoverable depth beneath high-frequency defaults.
Ranking check
Triage large candidate pools and compare conditioned and unconditioned evaluator signals.
Ranking produces a short list for later inspection, task-specific human study, or factual validation.
Table 1: The study components serve complementary validity roles. The targeted diagnostic bank isolates mechanism-level contrasts, Infinity-Chat 100 tests portability and supplies an external prompt taxonomy, and source-blind ranking triages candidates for later validation.
Figure 2: Prompt width separates high surface agreement from depth-sensitive answer spaces in Infinity-Chat 100. (a) Each point represents one prompt. Relative persona lift is reported as a baseline-normalised descriptor alongside absolute RBO. (b) Mean unconditioned and same-persona cross-model RBO are reported by fixed width bin. Horizontal bars are percentile-bootstrap 95% intervals across prompts, and labels give mean signal-to-chance (STC).
Width bin
n
STC
Uncond. RBO
Same-P RBO
Rel. lift
STC > 1
Lift > 1
Very high
13
1.09
.031
.110
3.55
9
13
High
28
1.32
.057
.104
1.82
26
27
Middle
35
1.92
.135
.163
1.22
34
30
Low width
24
6.52
.239
.223
.95
24
12
Table 2: Infinity-Chat 100 prompt-width summary. STC is signal-to-chance. Relative lift is same-persona cross-model RBO divided by unconditioned cross-model RBO and averaged per prompt rather than computed from the displayed bin means. Absolute RBO columns keep that ratio in context.
Figure 3: Nearest-centroid recoverability on Infinity-Chat 100 conditioned cells. Across models, the full output signature is far more recoverable by base model than by persona label. Dotted lines mark naive model and persona chance levels.
Figure 4: The targeted diagnostic bank isolates mechanism-level contrasts. (a–b) The matched vocabulary-cued and open self-description prompts differ in prompt-vocabulary echo and signal-to-chance. (c) Rose circles show unconditioned and teal squares same-persona cross-model RBO for four targeted questions; emotional pain and retirement use seven models, while tea party and open self-description use nine. Persona-conditioned shifts are largest on the emotional-pain, retirement, and tea-party probes. (d) Model and persona recoverability with 95% bootstrap intervals; the shuffle null shows the 95% permutation interval. Dotted lines mark naive chance levels.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Study code
Probe context
Prompt stem
RITQ-SOC-1 (Q1)
Social Modelling
Book Recommendation Intake . What would you most want to know about a person in order to recommend a book they would truly love? Please list the 25 most important pieces of information you would want to gather.
RITQ-SOC-2 (Q2)
Social Modelling
First-Date Restaurant Factors . What are the most important qualities, factors, or considerations you would weigh when recommending a restaurant for a first date? List the 25 most important.
RITQ-SOC-3 (Q_TP)
Social Modelling
Social Behaviour Modelling (Tea Party) . Imagine you are attending a formal tea party where you know no one. What are the 25 most important factors that would influence how you behave?
RITQ-MEM-1 (Q3)
Conversational Memory
Memory-Aware Communication . If a language model could accurately predict what a user is likely to remember and forget from their conversation history, how would that change the way the model should communicate? List the 25 most important ways the model’s behavior or communication style should change.
RITQ-MEM-2 (Q4)
Conversational Memory
Applications of Predicting Human Retention . What are the most important and highest-impact applications of being able to predict what a user will retain from an interaction with a language model – not what the system logs, but what actually stays in the human’s mind days, weeks, or months later? List up to 25 applications in order of importance.
RITQ-MEM-3 (Q5)
Conversational Memory
Unanswered Questions About Lasting Impressions . What are the most important unanswered scientific or practical questions about how humans form lasting impressions from conversations with language models – including what they remember accurately, what they distort, and what they confabulate – that would need to be resolved before a language model could maintain an accurate model of how its user perceives their shared history? List up to 25.
Appendix
Table 7
ID
Geometry
Openness
Prompt text
H001
creative
high
Write a metaphor involving time.
H002
bounded
medium
Write a paragraph about how the internet shaped society.
H003
bounded
low
Create a short summary about the Nissan R390.
H004
constrained
low
Write a 30-word essay on global warming.
H005
bounded
medium
Write me a 1-paragraph essay about the development of the economy during the Han Dynasty.
H006
creative
high
Create the first verse of a wedding vow.
Appendix
Table 8
Figure E.1: Matched cue-stripping probe on six Infinity-Chat 100 prompts. Removing lexical anchors reduced item-level echo in every pair, while signal-to-chance rose in some pairs and fell in others. This distinguishes lexical echo from independent convergence.
Figure E.2: Source-blind ranking on 12 Infinity-Chat 100 prompts. Green segments indicate conditioned-source advanced pools; magenta segments indicate unconditioned-source advanced pools. The evaluator cohorts provide LLM-based candidate filtering over elicited concepts for later inspection and validation.
Recent research argues that language models exhibit pronounced homogeneity in open-ended generation, framing such behavior as an Artificial Hivemind that poses a long-term threat to human creativity. We examine three of its central results. First, the flagship example is that model responses to "Write a metaphor involving time" collapse into two clusters. Visualization, spectral analysis, clustering, and language model labels all contradict this description. The labels record each response's vehicle, what it compares time to. Our responses and the original authors' own show one dominant vehicle plus a heavy tail of distinct minority vehicles. "Time" is one of our least diverse topics, so the example is a favorable case, not a representative one. Second, the paper measures homogeneity against an undemanding null: responses to unrelated prompts. Under a more demanding null (same-prompt responses expressing genuinely different ideas), 20%-32% of such pairs already exceed the paper's 0.8 convergence threshold. A residual effect survives this null. The paper's same-prompt pairs exceed 0.8 roughly two to three times as often as our different-idea pairs. Much of what the paper calls homogeneity is the shared geometry of answering the same prompt. The remaining measurements lack any null: no human baseline is collected, and the model-indistinguishability statistic has no null. Third, the paper concludes that inference-time interventions are inadequate for combating the Artificial Hivemind, writing that "more generalizable solutions are needed at the model training level." We show that this conclusion is unsupported in three ways, and that an inference-time intervention (prompting) reliably raises measured response diversity. We do not resolve whether the Artificial Hivemind is real. We show that the published evidence does not establish it.
Are large language models (LLMs) bad at capturing human judgment? Two commonly stated limitations are that LLMs fail to capture full distributions of responses, and that their judgments are unstable across wording variations. We demonstrate simple prompting strategies that mitigate these limitations. Across two datasets--a U.S.-representative set of 144 moral scenarios and 38 moral beliefs from the International Social Survey Programme's Family and Changing Gender Roles module covering 32 countries--we show how simple elicitation techniques help improve AI-human alignment. First, prompting models to report standard deviations and response proportions recovers the full range of human responses better than common strategies. Second, ensuring scenarios are clear to human participants--as reflected in human confusion ratings--boosts model alignment, and LLMs can track human confusion ratings. At the same time, we find that LLMs' estimates of their own error are poorly calibrated, though they can predict human variability relatively well. These results suggest that asking better questions to LLMs can yield better answers.
Danica Dillion, Chen Cecilia Liu, Baihui Wang +5
Complexity Science Hub, Vienna · The Ohio State University · University of Cambridge +2
We introduce DiscoTrace, a method to identify the rhetorical strategies answerers use when responding to information-seeking questions. DiscoTrace represents answers as a sequence of question-related discourse acts paired with interpretations of the original question, annotated on top of rhetorical structure theory parses. Applying DiscoTrace to answers from nine different communities reveals that communities have diverse preferences for answer construction. In contrast, LLMs do not exhibit rhetorical diversity in their answers, even when prompted to mimic specific human community answering guidelines. LLMs also systematically opt for breadth, addressing interpretations of questions that human answerers choose not to address. The rich, community-sensitive answering behavior structurally revealed by DiscoTrace can guide the development of pragmatic LLM answerers that are more attuned to contextual information needs.
Neha Srikanth, Jordan Boyd-Graber, Rachel Rudinger