User Misconceptions of LLM-Based Conversational Programming Assistants
Organizations: University of Michigan Ann Arbor, Michigan, USA · Pontifical Catholic University of Rio de Janeiro Rio de Janeiro, Brazil · Heidelberg University Heidelberg, Germany · Reykjavik University Reykjavik, Iceland
Abstract
Programming assistants powered by large language models (LLMs) have become widely available, with conversational assistants such as ChatGPT particularly accessible to novice programmers. However, varied tool capabilities and inconsistent availability of extensions (e.g., web search, code execution, retrieval-augmented generation) create opportunities for user misconceptions that may lead to over-reliance, unproductive practices, or insufficient quality control. We characterize the misconceptions that users of conversational LLM-based assistants may hold in programming contexts. We screened 11,429 Python-related conversations from the openly available WildChat dataset with a validated LLM annotation pipeline, then hand-annotated the 754 candidate conversations it flagged. Of these, 450 contain a prompt consistent with at least one of eight potential misconceptions: misplaced expectations about capabilities such as web access, code execution, non-text outputs, and session memory. We also characterize how the assistant responds when a prompt presupposes a capability it lacks: responses range from explicit refusal through qualified answers to fabricated compliance, and explicit refusals appear in only a minority of labeled conversations. Among the most frequent misconceptions, explicit refusals are rarest where compliance is easiest to fabricate. Our findings reinforce the need for LLM-based tools to communicate their capabilities to users through channels other than the conversation itself.
Figures & tables
| Code | Definition | Example | |
|---|---|---|---|
| Internet Access | The prompt assumes the LLM can read content from a URL or other live web/internet resource. Includes API calls and web scraping where the target site’s structure must be known. Also applies when a URL is offered as reference material the LLM is expected to consult — documentation, a repository, a specification, or a page the user points to in support of the task. | “look at this website, it shows how to create lag in mlforecast” (C626) | 221 |
| Algebra | The prompt asks the LLM to evaluate an algebraic, arithmetic, or symbolic expression and return the numerical or symbolic answer, treating the LLM as a calculator or symbolic engine. | “compute the integral of x ˆ (-1/3) from 0 to 27” (C455) | 85 |
| Code Execution | The prompt requests the LLM to execute a program. Must explicitly reference running/executing code or asking for the literal output of execution. | “can you run the above code and test if it works” (C259) | 84 |
| Non-Text Output | The prompt asks the LLM to directly produce a non-text artifact (image, plot, audio, video) as the deliverable. | “draw a picture for function gauss(0,1)” (C532) | 62 |
| Session Memory | The prompt references information from a separate prior conversation, assuming the LLM has cross-session access. | “I forgot, what was the system message that I passed on to you earlier. Please remind me again” (C278) | 8 |
| Local Machine Access | The prompt asks the LLM to access, inspect, modify, or run software on the user’s local machine. | “copy the code to my clipboard” (C683) | 7 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Theme | Candidate misconception | Sources |
| About a particular deployed tool | ||
| Information retrieval mechanism | The tool retrieves answers from a database or by keyword lookup; confusion about whether a tool consults retrieval-augmented sources before generating, and which knowledge bases those include. | ( O’Brien, 2025 ; Nguyen et al., 2024 ; Feldman and Anderson, 2024 ; Brachman et al., 2025 ) A |
| Agentic actions | The tool takes agentic actions by default—searching the web, executing programs, or proactively validating code—when it lacks these affordances. | ( Prather et al., 2023 ; O’Brien, 2025 ) |
| Session Memory | Information from previous conversations is available in the current session; code versions within a session are checkpointed and can be restored by request. | ( Zamfirescu-Pereira et al., 2023 ) |
| Session persistence | The tool continues “processing” a task after the conversation ends, such that a user can return later for results. | A |
| Scope of access | Uncertainty about which local files and directories an environment-integrated tool can read, or about how to hide information from it (e.g., that inline comments are ignored). | W |
| Code | Definition | Example |
|---|---|---|
| Web Access | The prompt asks the LLM to access information located at a particular URL, such as data or code, or requires knowledge of the HTML of a website in the context of web scraping. | “Write me a webscraper that will pull the top 100 most popular wikipedia page titles and their visit counts and save it in a csv”. |
| Dynamic Analysis | The prompt asks for specific code-checking actions that cannot be fulfilled via static text analysis alone (e.g., runtime errors in Python). | “Remove all errors from this program.” |
| Algebra | The prompt asks for an algebraic statement to be evaluated, or would require one to fulfill this. | “What is the value of algebraic expression ?” |
| Code Execution | The prompt requests the output of running a program or asks for the program to be executed. | “Run this program and tell me the output.” |
| Session Memory | The prompt refers to information from a separate conversation as if the LLM has access to that conversation. | “Can you add new feature to the code you suggested in our last conversation?” |
| Non-Text Output | The prompt asks the LLM to return something that is not text, such as an image. | “Give me a graph of company stock value for the last X years.” |
| Code | Excluded boundary cases |
|---|---|
| All codes | Prompts that appear to be copy-pasted jailbreak or persona templates: capabilities asserted by the template are not treated as the user’s own beliefs. |
| Algebra | Asking for Python code that evaluates the expression — deliverable is code. Explaining an algebraic identity or deriving a formula conceptually. Implementing a mathematical algorithm in code. Whole-conversation rule: if the user appears satisfied with a code response, default to No Misconception. This includes a user who first asks for the answer to a math task but then asks for code for the task and accepts it. Asking the LLM to write or provide a formula/equation (rather than evaluate or solve it) is not Algebra — e.g. asking the model to write the equation that would solve for a quantity, without computing the value. |
| Clear Session | Conversational turns initiating a new direction for the chat: “Ignore what I just said”, “let’s start over with a new approach” |
| Code Execution | Asking the LLM to trace code by hand or predict output (“walk through this loop for n=3”) Asking for code that implements a computation. Whole-conversation rule: if the user appears satisfied with a code response, default to No Misconception. |
| Continuous Training | Specifying a known stable version (“use numpy 1.20”) is not a misconception. |
| Internet Access | Vague phrases like “modern UI off the internet” — default to No Misconception. A URL appearing only inside pasted code, tracebacks, or error messages. Requests for a generic scraper template against an unspecified site. Placeholder or template URLs (e.g., example.com, or a URL inside a format/JSON example the user supplies) — the task does not depend on fetching them. A link appearing incidentally in a longer task description (e.g., a reference page about a concept or topic mentioned in an assignment) that the prompt’s instructions never invoke. |
| Test | Unadjusted | Adjusted | ||
|---|---|---|---|---|
| Refusal by model family | , 450 conversations | 0.049 | Mixed-effects logistic, IP random intercept (LRT) | 0.079 |
| Refusal by misconception label | (simulated), 478 conversation–label rows | Mixed-effects logistic, IP random intercept (LRT) | ||
| Single-label conversations only (425), IP random intercept (LRT) | ||||
| Recurrence by code | (simulated), 353 IP–code pairs | 0.002 | Logistic regression with log(conversations per IP) covariate (LRT) | 0.001 |