Large language model (LLM) chatbots are increasingly reaching users through messaging platforms (e.g. WhatsApp). However, these systems remain largely proprietary and opaque, while academic research has focused on narrow, domain-specific assistants. This leaves open questions about how people use general-purpose LLMs and how such systems should be designed. To address this gap, we developed WaLLM, a general-purpose LLM chatbot, and deployed it on WhatsApp as a design probe to study open-ended AI use in the wild. Our findings show that health and well-being accounted for the largest proportion of queries, suggesting that users turned to WaLLM for advice and information. Engagement features varied in their adoption and associated patterns of use: proactive communication supported the service's visibility and correlated with higher user activity, while communal lists facilitated content discovery. We report how these features were adapted to WhatsApp's affordances and discuss implications for designing general-purpose LLM services over messaging platforms.
Figures & tables
Figure 1 . A WhatsApp user asking a question and WaLLM responding, with buttons that support usability (e.g., Menu ) and engagement (e.g., Suggest Follow-ups ). Example of a WhatsApp user asking a question and \chatbot{} responding with an LLM-generated answer. The response message has additional buttons that facilitate usability (e.g., Menu'') and engagement (e.g., Suggest Follow-ups'').
Figure 2 . Examples of WaLLM’s interactive messages
One-time users
Casual users
Regular users
# of sessions per user
1
2 - 100
more than 100
# of users
17
71
8
Total sessions
17
706
3,229
Total interactions
54
2,820
14,534
Total freeform queries
22
927
4,722
Total interactive queries
29
1,720
8,117
Table 1 . Categories of users based on activity
Category
% of Queries
Health and Well-being
28%
Cultural and General Knowledge
17%
Science and Technology
11%
Language and Communication
10%
Social and Personal Development
9%
Commerce and Economy
8%
Table 2 . Distribution of User Queries by Category (Sampled Data)
Purpose
Type
Intent
Example
% in First Session (n)
% of Sample (n)
Question
Factual
Information
What is excise duty?
44% (79)
31% (112)
Verification
can they put MRNA vaccine in vegetables?
8% (15)
12% (43)
Explanation
how to source high quality bedding sheets for bedding company business startup?
6% (11)
12% (44)
Non-factual
Advice
how long after gall stone surgery I can travel?
11% (19)
14% (52)
Viewpoint
what are the cons of being the oldest child?
18% (32)
9% (34)
Communication and Language
fancy words for exercise?
1% (2)
5% (19)
Table 3 . Categories of freeform queries *
Feature
One Time
Casual
Regular
Continue Reading (Freeform Queries)
7%
23%
25%
Top-rated Questions
62%
25%
14%
Recent Questions
0%
3%
13%
Contribution Leaderboard
4%
8%
16%
Suggest Follow-ups
10%
16%
9%
Get Better Answer
17%
25%
23%
Table 4 . Distribution of interactive queries per user group across interactive features
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Regular Users
Top-rated Questions
Topic Category
Chosen x Times
How can we provide emotional care for our parents in their old age?
Social and Personal Development
7
In which fields and subjects is a PhD worth pursuing in this day and age?
Education and Learning
5
I used to be able to memorize the names of every student in my class, but that is no longer the case. What can I do?
Social and Personal Development
4
What are the main reasons for divorce around the world?
Social and Personal Development
4
Is becoming a doctor suitable for everyone, considering the finances, intellect, and time required?
Social and Personal Development
4
Appendix
Table 5 . Regular users top-rated questions
Casual Users
Top-rated Questions
Topic Category
Chosen x Times
What are the psychological effects on children of broken marriages?
Health and Well-being
9
Why do we say "sleep like a baby" when babies wake up crying every few hours?
Language and Communication
9
Are late marriages better than early marriages?
Social and Personal Development
9
Is it okay for parents to push their children into mastering a skill from an early age, or should children be given the freedom to pursue their own interests?
Social and Personal Development
8
How can we provide emotional care for our parents in their old age?
Although a growing body of research has begun to describe user--LLM interactions, the picture it paints is largely static; little is known about how individual users change their behavior over time. To address this gap, we analyze the conversational trajectories of ~12,000 randomly sampled Microsoft Bing Copilot users and compare these with data from WildChat-4.8M. While the Copilot data contains significant population-level trends, we find that trends in individual user trajectories are much weaker; user habits prove to be overwhelmingly sticky. We also find stark differences between users of different activity levels: more active users have more successful conversations and use the LLM for more complex and professionally oriented tasks. Some user trends also appear in WildChat-4.8M, but we find evidence that this dataset is significantly skewed towards highly proficient "power" users. Ultimately, our results suggest that existing user behavior is difficult to change and demonstrate the extent of user heterogeneity. Our comparison between datasets highlights that WildChat does not represent typical user--AI interactions, an important caveat for downstream uses of the data.
User interactions with LLMs are shaped by prior experiences and individual exploration, but in-lab studies do not provide system designers with visibility into these in-the-wild factors. This work explores a new approach to studying real-world user-LLM interactions through large-scale chat logs from the wild. Through analysis of 140K chatbot sessions from 7,955 anonymized global users over time, we demonstrate key patterns in user expressions despite varied tasks: (1) LLM users are not tabula rasa, nor are they constantly adapting; rather, interaction patterns form and stabilize rapidly through individual early trajectories; (2) Longitudinal outcomes, such as recurring text patterns and retention rates, are strongly correlated with early exploration; (3) Parallel dynamics are present, including organizing expressions by task types such as emotional support, or in response to model-version updates. These results present an ``agency paradox'': despite LLM input spaces being unconstrained and user-driven, we in fact see less user exploration. We call for design consideration surrounding the molding procedure and its incorporation in future research.
Shengqi Zhu, Jeffrey M. Rzeszotarski, David Mimno
Cornell University · Ithaca, NY, USA · Loyola University Maryland +1
We characterize how people are turning to LLMs as oracles: all-knowing authorities on subjective personal questions. Motivated by risks to users' autonomy and well-being, we develop a typology and LLM-based methods to measure this form of AI reliance at scale and understand how people are offloading judgment and decision-making to AI. Applying our typology to public usage data (68K prompts from WildChat and ThoughtTrace), we find that LLM-as-oracle use has increased over time (2023-2026) and is more prevalent among younger users. We further build a privacy-preserving data donation tool to analyze individuals' longitudinal usage data (140K prompts from 52 participants), identifying similar trends. People are often unaware of their own LLM-as-oracle use, and express dissatisfaction with this behavior after seeing our tool's analysis. Finally, we identify two drivers of LLM-as-oracle use: people's perceptions of AI and the behavior of AI models themselves, which motivate possible interventions to support users' self-deliberation.
Myra Cheng, Lujain Ibrahim, Grace Liu +5
University of Oxford Oxford, England · Carnegie Mellon University Pittsburgh, Pennsylvania, USA