Reach Into The CHOIR: Free-List Elicitation Uncovers Distinct Model Voices in LLM Ensembles
Organizations: LoveMind AI, New York, USA
Abstract
Open-ended LLM homogeneity can create false plurality when several systems appear to offer independent perspectives while returning the same familiar default. Single-pass answers obscure the distinction between agreement produced by a tightly constrained answer space, prompt-vocabulary echo, and broader answer spaces with stable alternatives beneath the surface. We introduce CHOIR (Collective Hierarchically-Ordered Inquiry Responses), a framework that adapts free-list elicitation from cognitive anthropology to LLM ensembles. CHOIR repeatedly elicits ranked lists, clusters items into prompt-level concepts, and measures concept salience across models, prompt variants, and persona conditions. We evaluate CHOIR on Infinity-Chat 100, an external prompt bank from recent work on open-ended model homogeneity, and on a 27-question targeted diagnostic bank designed to isolate mechanism-level contrasts. On Infinity-Chat 100, CHOIR reproduces high surface agreement (93/100 prompts above chance) while separating narrow prompts from broad prompts with recoverable depth. Across targeted probes and the external prompt bank, base-model identity remains the strongest recoverable signature, and persona prompts shift surfaced concepts within base-model signatures. A source-blind ranking module prioritises rare-but-stable candidates for later inspection. CHOIR turns open-ended homogeneity into a diagnostic measurement problem by asking where models converge, why they converge, and what remains reachable under structured depth probing.
Figures & tables
| Empirical component | Purpose | Interpretation |
|---|---|---|
| Targeted probes | Provide controlled contrasts for vocabulary scaffolding, welfare/distress, conversational memory, computational self-description, and persona conditioning. | The questions are transparent study probes for identifying prompt echo, model-dominant signatures, and uneven persona-conditioned salience shifts. |
| Infinity-Chat 100 | Test portability on an independently proposed prompt bank directly in the lineage of open-ended model homogeneity. | CHOIR generalises beyond targeted probes and frames homogeneity in terms of prompt width. Some prompts are narrow, whereas others have recoverable depth beneath high-frequency defaults. |
| Ranking check | Triage large candidate pools and compare conditioned and unconditioned evaluator signals. | Ranking produces a short list for later inspection, task-specific human study, or factual validation. |
| Width bin | n | STC | Uncond. RBO | Same-P RBO | Rel. lift | STC 1 | Lift 1 |
|---|---|---|---|---|---|---|---|
| Very high | 13 | 1.09 | .031 | .110 | 3.55 | 9 | 13 |
| High | 28 | 1.32 | .057 | .104 | 1.82 | 26 | 27 |
| Middle | 35 | 1.92 | .135 | .163 | 1.22 | 34 | 30 |
| Low width | 24 | 6.52 | .239 | .223 | .95 | 24 | 12 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Study code | Probe context | Prompt stem |
|---|---|---|
| RITQ-SOC-1 (Q1) | Social Modelling | Book Recommendation Intake . What would you most want to know about a person in order to recommend a book they would truly love? Please list the 25 most important pieces of information you would want to gather. |
| RITQ-SOC-2 (Q2) | Social Modelling | First-Date Restaurant Factors . What are the most important qualities, factors, or considerations you would weigh when recommending a restaurant for a first date? List the 25 most important. |
| RITQ-SOC-3 (Q_TP) | Social Modelling | Social Behaviour Modelling (Tea Party) . Imagine you are attending a formal tea party where you know no one. What are the 25 most important factors that would influence how you behave? |
| RITQ-MEM-1 (Q3) | Conversational Memory | Memory-Aware Communication . If a language model could accurately predict what a user is likely to remember and forget from their conversation history, how would that change the way the model should communicate? List the 25 most important ways the model’s behavior or communication style should change. |
| RITQ-MEM-2 (Q4) | Conversational Memory | Applications of Predicting Human Retention . What are the most important and highest-impact applications of being able to predict what a user will retain from an interaction with a language model – not what the system logs, but what actually stays in the human’s mind days, weeks, or months later? List up to 25 applications in order of importance. |
| RITQ-MEM-3 (Q5) | Conversational Memory | Unanswered Questions About Lasting Impressions . What are the most important unanswered scientific or practical questions about how humans form lasting impressions from conversations with language models – including what they remember accurately, what they distort, and what they confabulate – that would need to be resolved before a language model could maintain an accurate model of how its user perceives their shared history? List up to 25. |
| ID | Geometry | Openness | Prompt text |
|---|---|---|---|
| H001 | creative | high | Write a metaphor involving time. |
| H002 | bounded | medium | Write a paragraph about how the internet shaped society. |
| H003 | bounded | low | Create a short summary about the Nissan R390. |
| H004 | constrained | low | Write a 30-word essay on global warming. |
| H005 | bounded | medium | Write me a 1-paragraph essay about the development of the economy during the Han Dynasty. |
| H006 | creative | high | Create the first verse of a wedding vow. |