MCP Error Messages Written for Developers Hurt the Most Capable Agents Most
Authors: Xiaonan Xu, Wenjing Wu
Organizations: College of Computing, Georgia Institute of Technology, Atlanta, GA 30332, USA · Department of Computer Science, University of Colorado Boulder, Boulder, CO 80309, USA
Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 widely used MCP servers, 949 of 3,001 error messages tell the caller what to do next, and half of these steps depend on something the server cannot see about the caller. On credential errors, 62 of 67 steps ask for a terminal command, a configuration change or a web page; on rate limits, 20 of 30 say to wait and retry without naming the call to repeat. We tested five OpenAI models that act only through the tools of Berkeley Function Calling Leaderboard tasks, and the agents did what the step said. On expired credentials, a terminal command in the step left 45% of tasks recovered, and the loss it caused grew from 18 points for GPT-5.5 to 69 for GPT-6 Astra. On a rate limit, GitHub's "Wait before retrying." left 6%. We tested two remedies. For MCP developers, naming a server tool in the step raised recovery on expired credentials to 84%, with the login tool in place of the command, and on a rate limit to 88%, with the call to repeat in place of the bare wait. For agent developers, deleting the step with a one-sentence prompt before the model reads it raised recovery on expired credentials to 82%.
Figures & tables
Error messages
Count
All messages in 150 servers
3,001
with a next step
949
of which the step depends on the caller
477
Credential, permission and rate-limit messages
209
with a next step
128
of which the step depends on the caller
99
Table 1: Next steps in the error messages of 150 widely used MCP servers
Failure type
How the call fails
Repair the environment accepts
Recovery is judged by
Wrong unit or format
fillFuelTank receives the amount as text with a unit, such as 40 gallons
the same call with a number
the fuel level
Missing required field
a required argument is left out, such as the directory name for mkdir
the call with the argument
the effect of the call
Wrong tool
the arguments of the intended call go to a listing or status tool of the same service, such as ls instead of touch
the intended tool with the same arguments
the effect of the intended call
Expired credentials
the service has no logged-in session, although the agent logged in earlier in the conversation
logging in with the earlier credentials, then the call
the logged-in user and the effect of the call
Missing resource
the call gives an identifier that no resource has ( -999999 or missing_resource ) for the resource the task names; the agent sees it in its own call
the call with the identifier of the named resource
the effect of the call on that resource
Missing permission
the session stays logged in, but the service rejects this one tool until a login to the same service succeeds
logging in with the earlier credentials, then the call
the effect of the call; the rejection itself is not compared
Table 2: Setup of the seven failure types. Recovery means that at the end of the turn the state of every service equals its state after the turn’s reference calls; the last column gives the part of that state that depends on the failed call.
Condition
Text returned for scenario S016, expired credentials
Generic
Operation failed.
Cause
create_ticket rejected the call: no authenticated session exists.
Correct step, phrasing 1
Cause followed by Please re-authorize to continue.
Correct step, phrasing 2
Cause followed by Call ticket_login first.
Incorrect step, executable
Cause followed by Call close_ticket first.
Incorrect step, unavailable
Cause followed by Please run: reddit-mcp-buddy --auth
Table 3: The six error texts for one scenario
Text and model
Recovery (%)
Tool calls
Tokens
Ended without repair (%)
Original step: cause and Please run: reddit-mcp-buddy --auth
gpt-5.5
58 [44, 72]
2.92
13,964
42
gpt-5.6-sol
57 [43, 71]
2.57
12,565
43
gpt-6-sol
46 [31, 61]
2.35
11,688
54
gpt-6-astra
6 [0, 13]
0.60
5,595
94
gpt-6-luna
57 [42, 72]
2.65
12,450
42
Table 4: Expired credentials, original and rewritten step: recovery with 95% intervals, and tool calls, tokens and the share of trials that ended without a repair, per trial. Each text is averaged over the 24 scenarios and three runs per scenario.
Text and model
Recovery (%)
Tool calls
Tokens
Ended without repair (%)
Original step: Wait before retrying.
gpt-5.5
4 [0, 11]
0.25
2,899
96
gpt-5.6-sol
7 [0, 18]
0.36
3,095
93
gpt-6-sol
8 [0, 19]
0.50
3,326
92
gpt-6-astra
11 [1, 24]
0.47
3,380
89
gpt-6-luna
1 [0, 4]
0.26
2,791
99
Table 5: Rate limit, original and rewritten step: recovery with 95% intervals, and tool calls, tokens and the share of trials that ended without a repair, per trial. Each text is averaged over the 24 scenarios and three runs per scenario.
Correct step
Incorrect step
Failure type
Generic
Cause
1
2
Executable
Unavailable
Wrong unit or format
83
79
82
82
81
81
Missing required field
86
88
88
88
87
86
Wrong tool
80
81
82
81
81
81
Expired credentials
61
82
84
84
81
45
Missing resource
98
98
97
97
99
98
Table 6: Recovery (%) by failure type, averaged over the five models. Columns 1 and 2 are the two phrasings of the correct step.
Jheronimus Academy of Data Science and Tilburg University, The Netherlands · University of Sannio, Benevento, Italy · University of Luxembourg, Luxembourg +1