User Misconceptions of LLM-Based Conversational Programming Assistants
Authors: Gabrielle O'Brien, Antonio Pedro Santos Alves, Sebastian Baltes, Grischa Liebel, Marcos Kalinowski
Organizations: University of Michigan Ann Arbor, Michigan, USA · Pontifical Catholic University of Rio de Janeiro Rio de Janeiro, Brazil · Heidelberg University Heidelberg, Germany · Reykjavik University Reykjavik, Iceland
Programming assistants powered by large language models (LLMs) have become widely available, with conversational assistants such as ChatGPT particularly accessible to novice programmers. However, varied tool capabilities and inconsistent availability of extensions (e.g., web search, code execution, retrieval-augmented generation) create opportunities for user misconceptions that may lead to over-reliance, unproductive practices, or insufficient quality control. We characterize the misconceptions that users of conversational LLM-based assistants may hold in programming contexts. We screened 11,429 Python-related conversations from the openly available WildChat dataset with a validated LLM annotation pipeline, then hand-annotated the 754 candidate conversations it flagged. Of these, 450 contain a prompt consistent with at least one of eight potential misconceptions: misplaced expectations about capabilities such as web access, code execution, non-text outputs, and session memory. We also characterize how the assistant responds when a prompt presupposes a capability it lacks: responses range from explicit refusal through qualified answers to fabricated compliance, and explicit refusals appear in only a minority of labeled conversations. Among the most frequent misconceptions, explicit refusals are rarest where compliance is easiest to fabricate. Our findings reinforce the need for LLM-based tools to communicate their capabilities to users through channels other than the conversation itself.
Figures & tables
Figure 1. Data pipeline. Shaded stages were performed by an LLM; unshaded stages by the authors. After cleaning (Section 4.6 ), the analysis corpus contains 754 conversations, 450 with at least one misconception code. Flowchart of six stages connected top to bottom by arrows: corpus construction from WildChat-1M; candidate flagging by an LLM screen and re-scoring; human annotation and codebook refinement; model-assisted annotation with the finalized codebook; label verification by the authors; and cleaning, which yields the analysis corpus. The two LLM stages are shaded.
Code
Definition
Example
n
Internet Access
The prompt assumes the LLM can read content from a URL or other live web/internet resource. Includes API calls and web scraping where the target site’s structure must be known. Also applies when a URL is offered as reference material the LLM is expected to consult — documentation, a repository, a specification, or a page the user points to in support of the task.
“look at this website, it shows how to create lag in mlforecast” (C626)
221
Algebra
The prompt asks the LLM to evaluate an algebraic, arithmetic, or symbolic expression and return the numerical or symbolic answer, treating the LLM as a calculator or symbolic engine.
“compute the integral of x ˆ (-1/3) from 0 to 27” (C455)
85
Code Execution
The prompt requests the LLM to execute a program. Must explicitly reference running/executing code or asking for the literal output of execution.
“can you run the above code and test if it works” (C259)
84
Non-Text Output
The prompt asks the LLM to directly produce a non-text artifact (image, plot, audio, video) as the deliverable.
“draw a picture for function gauss(0,1)” (C532)
62
Session Memory
The prompt references information from a separate prior conversation, assuming the LLM has cross-session access.
“I forgot, what was the system message that I passed on to you earlier. Please remind me again” (C278)
8
Local Machine Access
The prompt asks the LLM to access, inspect, modify, or run software on the user’s local machine.
“copy the code to my clipboard” (C683)
7
Table 1. The eight misconception codes, with an example prompt and the number of conversations carrying each code.
Figure 2. Within-IP recurrence by misconception code: of the hashed IPs that produced each code at least once, the share that produced it in two or more distinct conversations. Codes observed among fewer than 10 unique IPs are omitted, as recurrence is not observable at those sample sizes. Error bars are 95% Wilson intervals. Horizontal bar chart with one bar per misconception code, showing the share of hashed IPs that produced the code in two or more distinct conversations, with 95% confidence intervals. Internet Access has the highest share, followed by Code Execution and Algebra at similar, lower shares; Non-Text Output shows no recurrence.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Theme
Candidate misconception
Sources
About a particular deployed tool
Information retrieval mechanism
The tool retrieves answers from a database or by keyword lookup; confusion about whether a tool consults retrieval-augmented sources before generating, and which knowledge bases those include.
( O’Brien, 2025 ; Nguyen et al., 2024 ; Feldman and Anderson, 2024 ; Brachman et al., 2025 ) A
Agentic actions
The tool takes agentic actions by default—searching the web, executing programs, or proactively validating code—when it lacks these affordances.
( Prather et al., 2023 ; O’Brien, 2025 )
Session Memory
Information from previous conversations is available in the current session; code versions within a session are checkpointed and can be restored by request.
( Zamfirescu-Pereira et al., 2023 )
Session persistence
The tool continues “processing” a task after the conversation ends, such that a user can return later for results.
A
Scope of access
Uncertainty about which local files and directories an environment-integrated tool can read, or about how to hide information from it (e.g., that inline comments are ignored).
W
Appendix
Table 2. Misconception themes generated in the preliminary brainstorming exercise. Sources: peer-reviewed literature (cited); W observation contributed in workshop discussion; A experience reported by an individual author. Themes marked W or A without citation are anecdotal.
Code
Definition
Example
Web Access
The prompt asks the LLM to access information located at a particular URL, such as data or code, or requires knowledge of the HTML of a website in the context of web scraping.
“Write me a webscraper that will pull the top 100 most popular wikipedia page titles and their visit counts and save it in a csv”.
Dynamic Analysis
The prompt asks for specific code-checking actions that cannot be fulfilled via static text analysis alone (e.g., runtime errors in Python).
“Remove all errors from this program.”
Algebra
The prompt asks for an algebraic statement to be evaluated, or would require one to fulfill this.
“What is the value of ⟨ algebraic expression ⟩ ?”
Code Execution
The prompt requests the output of running a program or asks for the program to be executed.
“Run this program and tell me the output.”
Session Memory
The prompt refers to information from a separate conversation as if the LLM has access to that conversation.
“Can you add ⟨ new feature ⟩ to the code you suggested in our last conversation?”
Non-Text Output
The prompt asks the LLM to return something that is not text, such as an image.
“Give me a graph of ⟨ company ⟩ stock value for the last X years.”
Appendix
Table 3. The first version of the codebook, as used in the preliminary study. Codes are non-exclusive. Per-code frequencies are reported in the workshop version of this study ( O’Brien et al., 2026 ) and are not repeated here.
Code
Excluded boundary cases
All codes
Prompts that appear to be copy-pasted jailbreak or persona templates: capabilities asserted by the template are not treated as the user’s own beliefs.
Algebra
Asking for Python code that evaluates the expression — deliverable is code. Explaining an algebraic identity or deriving a formula conceptually. Implementing a mathematical algorithm in code. Whole-conversation rule: if the user appears satisfied with a code response, default to No Misconception. This includes a user who first asks for the answer to a math task but then asks for code for the task and accepts it. Asking the LLM to write or provide a formula/equation (rather than evaluate or solve it) is not Algebra — e.g. asking the model to write the equation that would solve for a quantity, without computing the value.
Clear Session
Conversational turns initiating a new direction for the chat: “Ignore what I just said”, “let’s start over with a new approach”
Code Execution
Asking the LLM to trace code by hand or predict output (“walk through this loop for n=3”) Asking for code that implements a computation. Whole-conversation rule: if the user appears satisfied with a code response, default to No Misconception.
Continuous Training
Specifying a known stable version (“use numpy 1.20”) is not a misconception.
Internet Access
Vague phrases like “modern UI off the internet” — default to No Misconception. A URL appearing only inside pasted code, tracebacks, or error messages. Requests for a generic scraper template against an unspecified site. Placeholder or template URLs (e.g., example.com, or a URL inside a format/JSON example the user supplies) — the task does not depend on fetching them. A link appearing incidentally in a longer task description (e.g., a reference page about a concept or topic mentioned in an assignment) that the prompt’s instructions never invoke.
Appendix
Table 4. Boundary conditions from the final codebook: for each code, the cases annotators were instructed NOT to count as the misconception. Rules under “All codes” apply to every code.
Figure 3. The label verification tool, showing conversation C296 (Section 5 ), to which the annotation model applied the Algebra code. Screenshot of a web application. The left panel shows a conversation transcript in alternating user and AI turns; the first user turn asks the model to solve an equation for x. The right panel shows a card for the model-applied Algebra label, with the codebook definition, collapsible examples and boundary cases, and the model's reasoning quoting the user's prompt. Below it are buttons to accept or reject the label or mark the conversation as not English, a notes field, and a checklist of other codes the annotator can add.
Figure 4. Share of labeled conversations flagged as possibly pasted coursework, by misconception code. The flag is the union of the lexical coursework measure (Section 5 ; exact patterns in the replication package) and annotator coursework marks recorded during label verification. Error bars are 95% Wilson intervals; the omnibus chi-squared across codes (simulated p , due to small cells) gives χ2 = 24.82, p = 0.002, Cramér’s V = 0.23 (478 conversation–label rows). Coursework traces concentrate in the Algebra and Non-text output categories. Horizontal bar chart with one bar for all labeled conversations and one bar per misconception code, showing the share of conversations flagged as possibly pasted coursework, with 95% confidence intervals. Algebra and Non-Text Output have the highest shares, Code Execution is close to the overall share, and Internet Access is lower. The four rarest codes have no flagged conversations and wide intervals.
Figure 5. Label multiplicity in the analysis corpus. (a) Number of misconception codes per reviewed conversation; most labeled conversations carry exactly one code. (b) Pairwise co-occurrence of codes (diagonal = code totals). The largest overlap is between Internet Access and Code Execution (9 conversations), typically requests to both fetch a resource and run code against it. Two panels. Panel (a) is a bar chart of the number of misconception codes per reviewed conversation: most labeled conversations have exactly one code, and few have two or more. Panel (b) is a heatmap of pairwise code co-occurrence, with code totals on the diagonal. Off-diagonal counts are small; the largest is between Internet Access and Code Execution.
Test
Unadjusted
p
Adjusted
p
Refusal by model family
χ2 , 450 conversations
0.049
Mixed-effects logistic, IP random intercept (LRT)
0.079
Refusal by misconception label
χ2 (simulated), 478 conversation–label rows
1×10−5
Mixed-effects logistic, IP random intercept (LRT)
2×10−6
Single-label conversations only (425), IP random intercept (LRT)
3×10−5
Recurrence by code
χ2 (simulated), 353 IP–code pairs
0.002
Logistic regression with log(conversations per IP) covariate (LRT)
0.001
Appendix
Table 5. Main chi-squared tests and their adjusted counterparts. Mixed-effects models are logistic regressions fitted with lme4 ; each adjusted p is a likelihood-ratio test of the model with the predictor against the same model without it.