Human--LLM interaction datasets shape our understanding of AI use and provide a foundation for downstream research, including training and evaluation of user models. In recent years, a growing number of datasets have sought to capture a representative picture of human--LLM interactions. But how different are the pictures these datasets provide, and what do those differences mean for research built on them? We study these questions across seven conversation datasets, spanning in-the-wild chat logs and human preference data. We begin by revisiting the dataset classification experiment of Torralba & Efros and find that neural network classifiers identify the source of a conversation from user messages alone well above chance, indicating distinctive dataset signatures. This separability persists after matching datasets on the dimensions of human-designed taxonomies, implying subtle differences that these taxonomies do not capture. We then examine the implications for user modeling: how dataset signatures propagate to the outputs of user models trained on these datasets; how dataset choice influences evaluations of user model quality and subsequent evaluations of LLM assistants paired with these user models; and how dataset classifiers can guide data selection for training user models. While each dataset is meant to capture a slice of 'real-world' interactions, our findings reveal the extent to which these slices diverge, and the consequences of those differences for research built on these foundations.
Figures & tables
Figure 1: Can you identify the source dataset? Nine example conversations, three each from WildChat-1M ( Zhao et al., 2024 ) , LMSYS-Chat-1M ( Zheng et al., 2024 ) , and HH-RLHF ( Bai et al., 2022 ) . On this three-way task, a neural classifier identifies the source with 83% overall accuracy and assigns over 90% probability to the correct source for each example shown here. Answers: WC-1M: 2, 6, 8; LMSYS: 3, 4, 9; HH-RLHF: 1, 5, 7.
Dataset
Description
Markers (prevalence)
WC-1M ( Zhao et al., 2024 )
WildChat-1M: voluntary interactions with free GPT-based services on HuggingFace Spaces.
PII redaction placeholder (0.8%)
WC-4.8M ( Zhao et al., 2024 )
An expanded release of WC-1M with newer assistant models, such as OpenAI o1.
PII redaction placeholder (1.3%)
LMSYS ( Zheng et al., 2024 )
LMSYS-Chat-1M: voluntary interactions on the Vicuna demo and Chatbot Arena platform.
NAME_ * (28.2%); [your answer] (1.1%)
Arena-human-preference-140K
Conversations reconstructed from released preference-vote histories during 2025.
-
ShareChat ( Yan et al., 2025 )
Publicly shared conversations retrieved from chatbot platforms.
<REDACTED> , <URL> , <DATE_TIME> (50.2%)
ShareGPT
User-shared ChatGPT history from ShareGPT.com. We use ShareGPT_Vicuna_unfiltered snapshot.
-
Table 1: Datasets used in our experiments. Please refer to Appendix A for filtering and sampling details. Marker prevalence is the percentage of conversations containing dataset-specific markers.
Dataset ↓ Setting →
Binary
Three-way
Incremental
WC-1M
✓
–
–
✓
✓
✓
–
–
✓
✓
✓
✓
✓
WC-4.8M
✓
–
–
–
–
✓
✓
✓
✓
✓
✓
✓
✓
LMSYS
–
✓
–
✓
✓
–
✓
✓
✓
✓
✓
✓
✓
Arena
–
✓
–
–
–
–
–
–
–
✓
✓
✓
✓
ShareChat
–
–
✓
✓
–
–
–
✓
–
–
✓
✓
✓
ShareGPT
–
–
✓
–
–
✓
–
–
–
–
–
✓
✓
Table 2: Dataset classification shows test accuracy well above chance. For each dataset combination, we train a multinomial logistic regression head that takes Qwen2.5-7B representations as an input.
Figure 4
Figure 3: Accuracy on WC-1M + LMSYS + ShareChat across training sample sizes and model sizes.
Accuracy (%)
Pair
Original
F
F+T
F+T+M
F+T+M+S
Δ [pp]
WC-1M / LMSYS
78.7
77.3
75.4
75.4
74.9
− 3.8
WC-1M / ShareGPT
86.2
83.5
82.9
83.8
83.2
− 3.0
LMSYS / ShareGPT
83.4
80.3
79.9
77.2
77.4
− 6.0
Table 4: Classification accuracy (%) after matching category distributions (chance is 50%). From left to right, columns add matched facets cumulatively: F (Function), T (Topic), M (Multi-turn), and S (Structural: numbers of turns and user tokens). Δ is the change from Original to F+T+M+S.
Figure 4: Category differences before and after matching on classifier scores. For each dataset pair and taxonomy facet, hollow and filled circles show the TVD between category distributions before and after matching; the bar marks the noise floor, the median TVD between two random halves of the same dataset across 1,000 bootstrap samples. Matching moves distances toward the noise floor.
Dataset pair
Real → real
Intent source
Real → synth.
Synth. → synth.
WC-1M / WC-4.8M
62.8
HH-RLHF
59.3
62.9
WC-4.8M / LMSYS
85.1
WC-1M
67.4
74.5
WC-4.8M / ShareChat
90.3
WC-1M
73.7
79.3
LMSYS / ShareChat
88.7
WC-4.8M
70.0
75.8
LMSYS / ShareGPT
83.4
WC-4.8M
63.6
72.8
ShareChat / ShareGPT
86.8
LMSYS
69.7
74.6
Table 5: Dataset classification accuracy (%) on synthetic conversations with Llama-based user models (chance is 50%). For each pair, two user models are conditioned on the same intents from the intent source. Real → real is a reference accuracy, training and testing on real conversations. Real → synth.: train on real, test on synthetic. Synth. → synth.: train and test on synthetic.
GSM8K
HumanEval
User model
Qwen
Llama
Δ
Qwen
Llama
Δ
WC-1M
9.6 ± 2.6
10.0 ± 3.1
-0.4 ± 2.6
14.2 ± 3.2
8.4 ± 2.6
5.8 ± 3.1
WC-4.8M
16.2 ± 4.1
13.8 ± 3.4
2.4 ± 3.9
18.0 ± 3.8
11.0 ± 3.2
7.0 ± 2.6
LMSYS
18.2 ± 4.2
17.6 ± 3.5
0.6 ± 3.7
18.0 ± 4.2
14.0 ± 3.7
4.0 ± 3.8
ShareChat
26.8 ± 4.4
22.0 ± 4.5
4.8 ± 4.5
17.0 ± 4.0
12.0 ± 3.2
5.0 ± 3.1
ShareGPT
24.8 ± 3.7
14.4 ± 2.9
10.4 ± 4.0
21.2 ± 4.6
9.8 ± 3.1
11.4 ± 4.3
Table 6: Task success rate (%) of Qwen2.5-14B-Instruct (Qwen) and Llama-3.1-8B-Instruct (Llama) as assistants when interacting with user models trained on different datasets (rows). Each entry averages 500 simulated conversations (100 problems × 5). 90% CI ( ± ) is calculated by first averaging the five outcomes for each problem, then using a Student’s t interval across the problem-level averages.
Train ↓ Test →
WC-1M
WC-4.8M
LMSYS
Arena
ShareChat
ShareGPT
HH-RLHF
WC-1M
6.41
6.70
6.14
6.15
8.38
6.02
5.39
WC-4.8M
6.53
6.36
6.17
6.06
8.28
6.07
5.46
LMSYS
6.61
6.80
5.43
6.17
8.31
6.03
5.19
Arena
6.69
6.82
6.22
5.75
8.29
6.14
5.29
ShareChat
6.71
6.85
6.22
6.13
7.33
6.09
5.30
ShareGPT
6.65
6.92
6.21
6.26
8.40
5.62
5.28
Table 7: Cross-dataset evaluation. Rows indicate training and columns indicate test (target) datasets. Each entry represents mean per-token perplexity measured on the test split. Bold and underlined entries indicate the lowest and second-lowest perplexity in each column, respectively.
Figure 5: Symmetric transfer gap ΔAB correlates with the A -to- B dataset classification accuracy.
Figure 6: Perplexity reduction relative to uniform random sampling at 50% of the training budget, PPL(50% random) - PPL(condition), for Qwen2.5-7B-Instruct user models (higher is better). The number under each dataset is the 50%-random perplexity. Classifier-guided selection outperforms random sampling at every matched budget. Llama-3-8B results are in Figure 8 .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Input
English
Eligible
Retained
HH-RLHF
169,352
168,805
145,740
50,000
ShareGPT
76,920
54,352
50,790
50,000
LMSYS-Chat-1M
1,000,000
766,545
415,087
50,000
Arena 140K (2025)
135,634
60,810
57,184
50,000
WildChat-1M
837,989
473,246
285,771
50,000
WildChat-4.8M
3,199,860
1,672,053
225,212
50,000
Appendix
Table 8: Curation and sampling counts for the snapshots used in this study. Input counts are raw records. English counts follow conversation reconstruction, structural validation, and language filtering. Eligible counts additionally reflect duplicate/template filtering and cross-source exclusions.
Figure 7: Percentage of conversations containing each annotated category in the original held-out WC-1M, LMSYS, and ShareGPT cohorts. A conversation can contain several categories, so columns need not sum to 100%. Each factor uses only conversations with complete accepted annotations.
Model size
Frozen
Fine-tuned
Δ [%]
0.5B
69.7
80.3
+10.6
1.5B
70.1
81.4
+11.3
3B
72.5
80.2
+7.7
7B
72.5
79.3
+6.8
Appendix
Table 9: Frozen language models vs. co-trained classifiers on WC-1M + LMSYS + ShareChat.
Real → synth.
Synth. → synth.
Dataset pair
Intent source
Qwen3.5
Llama3.1
Qwen2.5
Qwen3.5
Llama3.1
Qwen2.5
WC-1M / WC-4.8M
HH-RLHF
59.3
58.8
59.2
62.9
62.6
64.2
WC-4.8M / LMSYS
WC-1M
67.4
67.5
66.9
74.5
72.9
74.3
WC-4.8M / ShareChat
WC-1M
73.7
73.9
74.0
79.3
78.8
78.4
LMSYS / ShareChat
WC-4.8M
70.0
70.6
70.3
75.8
76.4
77.3
LMSYS / ShareGPT
WC-4.8M
63.6
64.8
65.2
72.8
70.6
70.0
Appendix
Table 10: User-model classification accuracy (%) on synthetic conversations with three different assistants (chance is 50%). For each pair, two user models are conditioned on the same intents from the intent source and interact with the same assistant; the Qwen3.5-9B columns repeat Table 5 . Real → synth.: train on real, test on synthetic. Synth. → synth.: train and test on synthetic.
User turns only
Assistant
Comparison
Chance
All
Single-turn
Multi-turn
turns only
Qwen3.5 vs. Llama3.1
50.0
53.4 (51.0 – 57.5)
50.5 (49.2 – 52.4)
54.6 (50.5 – 62.7)
99.2 (98.3 – 99.7)
Qwen3.5 vs. Qwen2.5
50.0
54.1 (50.7 – 59.6)
50.7 (49.4 – 52.0)
55.4 (50.9 – 63.8)
98.8 (97.8 – 99.6)
Qwen2.5 vs. Llama3.1
50.0
50.8 (49.8 – 53.0)
50.1 (49.4 – 51.1)
51.3 (50.0 – 54.9)
98.8 (97.8 – 99.6)
Three-way
33.3
35.9 (33.6 – 38.8)
33.7 (32.9 – 34.6)
36.6 (33.6 – 42.1)
98.3 (96.7 – 99.4)
Appendix
Table 11: Assistant classification accuracy (%) on synthetic conversations with the same intents and user model. Entries are means over 17 combinations of intent source and user model, with the range in parentheses. The first three columns use user turns only, with assistant responses masked; single-turn conversations have one user message and multi-turn conversations two or more. The last column uses assistant responses only, with user turns masked.
Train ↓ Test →
WC-1M
WC-4.8M
LMSYS
Arena
ShareChat
ShareGPT
HH-RLHF
WC-1M
6.60
7.05
6.72
6.65
8.89
6.70
5.78
WC-4.8M
6.78
6.47
6.81
6.54
8.75
6.77
5.84
LMSYS
6.93
7.30
5.60
6.71
8.90
6.72
5.60
Arena
7.03
7.33
6.91
6.10
8.93
6.93
5.76
ShareChat
7.10
7.36
6.94
6.71
7.70
6.85
5.83
ShareGPT
7.04
7.44
6.79
6.82
8.90
5.93
5.62
Appendix
Table 12: Cross-dataset evaluation. Rows indicate training and columns indicate test (target) datasets. Each entry represents mean per-token perplexity measured on the test split. Bold and underlined entries indicate the lowest and second-lowest perplexity in each column, respectively.
Figure 8: Perplexity reduction relative to uniform random sampling at 50% of the training budget, PPL(50% random) - PPL(condition), for Llama-3-8B user models (higher is better). The number under each dataset is the 50%-random perplexity. Classifier-guided selection outperforms random sampling at every matched budget.
Figure 9: Magnitude of classifier-score distribution change versus downstream NLL improvement. The horizontal axis shows Equation 4 , computed from trained classifier scores for target, randomly selected donor, and classifier-selected donor. The vertical axis shows target NLL reduction, ΔNLL .