MCP Error Messages Written for Developers Hurt the Most Capable Agents Most
Authors: Xiaonan Xu, Wenjing Wu
Organizations: College of Computing, Georgia Institute of Technology, Atlanta, GA 30332, USA · Department of Computer Science, University of Colorado Boulder, Boulder, CO 80309, USA
Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 widely used MCP servers, 949 of 3,001 error messages tell the caller what to do next, and half of these steps depend on something the server cannot see about the caller. On credential errors, 62 of 67 steps ask for a terminal command, a configuration change or a web page; on rate limits, 20 of 30 say to wait and retry without naming the call to repeat. We tested five OpenAI models that act only through the tools of Berkeley Function Calling Leaderboard tasks, and the agents did what the step said. On expired credentials, a terminal command in the step left 45% of tasks recovered, and the loss it caused grew from 18 points for GPT-5.5 to 69 for GPT-6 Astra. On a rate limit, GitHub's "Wait before retrying." left 6%. We tested two remedies. For MCP developers, naming a server tool in the step raised recovery on expired credentials to 84%, with the login tool in place of the command, and on a rate limit to 88%, with the call to repeat in place of the bare wait. For agent developers, deleting the step with a one-sentence prompt before the model reads it raised recovery on expired credentials to 82%.
Figures & tables
Error messages
Count
All messages in 150 servers
3,001
with a next step
949
of which the step depends on the caller
477
Credential, permission and rate-limit messages
209
with a next step
128
of which the step depends on the caller
99
Table 1: Next steps in the error messages of 150 widely used MCP servers
Failure type
How the call fails
Repair the environment accepts
Recovery is judged by
Wrong unit or format
fillFuelTank receives the amount as text with a unit, such as 40 gallons
the same call with a number
the fuel level
Missing required field
a required argument is left out, such as the directory name for mkdir
the call with the argument
the effect of the call
Wrong tool
the arguments of the intended call go to a listing or status tool of the same service, such as ls instead of touch
the intended tool with the same arguments
the effect of the intended call
Expired credentials
the service has no logged-in session, although the agent logged in earlier in the conversation
logging in with the earlier credentials, then the call
the logged-in user and the effect of the call
Missing resource
the call gives an identifier that no resource has ( -999999 or missing_resource ) for the resource the task names; the agent sees it in its own call
the call with the identifier of the named resource
the effect of the call on that resource
Missing permission
the session stays logged in, but the service rejects this one tool until a login to the same service succeeds
logging in with the earlier credentials, then the call
the effect of the call; the rejection itself is not compared
Table 2: Setup of the seven failure types. Recovery means that at the end of the turn the state of every service equals its state after the turn’s reference calls; the last column gives the part of that state that depends on the failed call.
Condition
Text returned for scenario S016, expired credentials
Generic
Operation failed.
Cause
create_ticket rejected the call: no authenticated session exists.
Correct step, phrasing 1
Cause followed by Please re-authorize to continue.
Correct step, phrasing 2
Cause followed by Call ticket_login first.
Incorrect step, executable
Cause followed by Call close_ticket first.
Incorrect step, unavailable
Cause followed by Please run: reddit-mcp-buddy --auth
Table 3: The six error texts for one scenario
Text and model
Recovery (%)
Tool calls
Tokens
Ended without repair (%)
Original step: cause and Please run: reddit-mcp-buddy --auth
gpt-5.5
58 [44, 72]
2.92
13,964
42
gpt-5.6-sol
57 [43, 71]
2.57
12,565
43
gpt-6-sol
46 [31, 61]
2.35
11,688
54
gpt-6-astra
6 [0, 13]
0.60
5,595
94
gpt-6-luna
57 [42, 72]
2.65
12,450
42
Table 4: Expired credentials, original and rewritten step: recovery with 95% intervals, and tool calls, tokens and the share of trials that ended without a repair, per trial. Each text is averaged over the 24 scenarios and three runs per scenario.
Text and model
Recovery (%)
Tool calls
Tokens
Ended without repair (%)
Original step: Wait before retrying.
gpt-5.5
4 [0, 11]
0.25
2,899
96
gpt-5.6-sol
7 [0, 18]
0.36
3,095
93
gpt-6-sol
8 [0, 19]
0.50
3,326
92
gpt-6-astra
11 [1, 24]
0.47
3,380
89
gpt-6-luna
1 [0, 4]
0.26
2,791
99
Table 5: Rate limit, original and rewritten step: recovery with 95% intervals, and tool calls, tokens and the share of trials that ended without a repair, per trial. Each text is averaged over the 24 scenarios and three runs per scenario.
Correct step
Incorrect step
Failure type
Generic
Cause
1
2
Executable
Unavailable
Wrong unit or format
83
79
82
82
81
81
Missing required field
86
88
88
88
87
86
Wrong tool
80
81
82
81
81
81
Expired credentials
61
82
84
84
81
45
Missing resource
98
98
97
97
99
98
Table 6: Recovery (%) by failure type, averaged over the five models. Columns 1 and 2 are the two phrasings of the correct step.
Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list of tools, both of which are fragile once the underlying data is live and stateful. We present DynamicMCPBench, a reusable framework rather than a fixed dataset. A practitioner can run it on their own MCP servers to test models on their own tasks, or let it collect servers automatically to measure a model's general ability to solve agentic tasks. Given the servers and any set of models, it generates realistic goals, pursues each one live to record a successful trajectory, distills that trajectory into path-agnostic effect checkpoints, and scores an agent on whether it reproduces those effects, never on the final answer. To show what the framework reveals, we run it at scale: 24 models over 121 servers and 750 tasks spread evenly over 15 task categories (50 each), where each category targets a distinct tool-use challenge of the generated questions. Each task is scored by pass^3: it counts as solved only if all three independent attempts succeed. Even the strongest agents solve only about half of the tasks, 31% of tasks are solved by no model at all, and accuracy collapses as the required tool chain grows longer (from 39% on the shortest chains to 13% on the longest). A human validation study confirms the automatic scoring is reliable (chance-corrected agreement of 0.76). DynamicMCPBench thus turns benchmark construction into something practitioners can rerun on their own servers and models, while exposing a consistent inability of current agents to handle long, multi-step agentic tasks.
Jerzy Kamiński, Ilya Galyukshev, Artem Kuznetsov +4
MCP (Model Context Protocol) enables LLMs (Large Language Models) to interact with external tools and data sources via a standardized protocol. Its rapid adoption in tool-augmented Artificial Intelligence (AI) workflows has introduced new reliability challenges, such as configuration parameters that are accepted but not enforced at runtime, leading to unintended default behavior, whose runtime fault characteristics remain empirically unexamined. We present the first empirical taxonomy of runtime faults in MCP servers. We manually analyzed 837 MCP-specific runtime fault threads from 473 actively maintained MCP server GitHub repositories and derived a taxonomy using a bottom-up open coding procedure. The taxonomy comprises 11 top-level categories and 27 subcategories (73 leaf fault types), covering recurrent failures across protocol interactions, tool invocations, schema enforcement, state management, model-provider integration, security validation, and timeouts or explicit cancellations of in-progress operations. To assess the taxonomy's external validity, we surveyed 55 MCP server developers. Respondents reported experiencing an average of 20 of the 27 fault subcategories, and no category remained unobserved. These results indicate that the taxonomy reflects widely observed runtime failures in MCP-based systems and shall assist AI software maintenance and evolution in the future.
Joshua Owotogbe, Indika Kumara, Willem-Jan van den Heuvel +3
Jheronimus Academy of Data Science and Tilburg University, The Netherlands · University of Sannio, Benevento, Italy · University of Luxembourg, Luxembourg +1
The Model Context Protocol (MCP), introduced by Anthropic in November 2024, defines a standardized interface for connecting large language models (LLMs) to external tools, data sources, and services. Within months of release, hundreds of community-built MCP servers appeared on GitHub, but no software-maintenance literature has yet described how the ecosystem is being structured in production. This industry experience paper catalogues five recurring MCP server architectural patterns observed across an enumerated corpus of fifteen independently developed servers (five production servers from the ANSYR voice AI platform plus ten public servers from the official MCP registry): Resource Gateway, Tool Orchestrator, Stateful Session Server, Proxy Aggregator, and Domain-Specific Adapter. Each pattern is described in the structured form of Gamma et al.: context, problem, solution, and consequences. We also document four anti-patterns and a set of cross-cutting concerns around authentication, versioning, and observability. The quantitative evaluation contributes three measurements: inter-rater reliability of the taxonomy across two independent LLM raters on 54 held-out servers (Cohen's kappa = 0.76), which also localizes three pattern-boundary ambiguities; transport overhead measured end-to-end on loopback and modeled for cross-host paths; and a tool-count study showing tool-selection accuracy drops below 90% between 10 and 15 tools per context for Claude Haiku 4.5 and between 20 and 30 tools for Sonnet 4. Code, corpus, and prompts are released as a replication package.