Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step. When an agent is the executor that predictability is normally lost: the prescribed procedure goes into the system prompt, and only a final answer comes back. We deliver the procedure step by step over the Model Context Protocol (MCP) instead: a server releases one step at a time, the agent executes it, and each step returns a structured step_output. This trades autonomy for predictability, and two properties then follow by construction, independent of the executor. The execution path is prescribed before the run, so the process is predictable in advance rather than reconstructed afterwards; and the completed step records form a machine-readable execution log that downstream tooling can audit and optimize step by step. Evaluating 15,475 trials across 13 SOP-Bench domains and four open-weight executors from frontier (Kimi K2.5) to lightweight (Ministral 3 8B), we find step-level delivery makes the executed process predictable and inspectable for every executor, and additionally raises accuracy when the executor is small. Across all four, process adherence rises significantly (76-95% to 95-99%) and ungrounded answers (correct outputs produced without executing the SOP) near-vanish, falling from 2.1-4.5% to 0.2-0.3% of trials (all 95% CIs exclude zero); under prompt-based delivery, 31-49% of correct answers on know_your_business bypass the SOP entirely, even for the frontier executor. Accuracy is where the executor's capability enters: the lightweight executor gains +6.5pp grounded accuracy because supplying the process externally removes a reconstruction burden it cannot carry, while capable ones trade a small raw-accuracy decrement for a predictable, auditable process.
Figures & tables
Executor
Version
Tier
max_tok
Ministral
3 8B
light
4,096
DeepSeek
V3.2
large
8,192
GLM
5
large
8,192
Kimi
K2.5
frontier
8,192
Table 1: Executors: family name (used throughout), version, capability tier, and output-token cap. Bedrock model IDs: mistral.ministral-3-8b-instruct , deepseek.v3.2 , zai.glm-5 , moonshotai.kimi-k2.5 .
Condition
TSR
Grounded
Δ
Ungr%
plain_sop
60.8%
60.6%
+0.2pp
0.3%
plain_enriched
61.5%
61.2%
+0.2pp
0.4%
rfc2119_sop
82.6%
78.4%
+4.2pp
5.1%
sop_mcp
78.0%
77.8%
+0.1pp
0.2%
Table 2: Bridge study (DeepSeek V3.2, 12 domains, domain-mean). Grounded TSR counts only correct answers from adherent trials. Δ = accuracy from shortcuts. Ungr% = fraction of correct answers that are ungrounded.
Adherence
Consistency
Executor
RFC
MCP
Δ
RFC
MCP
Δ
Ministral
76%
95%
+19 *
53%
58%
+5
DeepSeek
86%
97%
+11 *
65%
78%
+14 *
GLM
92%
99%
+7 *
70%
77%
+7 *
Kimi
95%
99%
+4 *
74%
79%
+4
Table 3: Adherence (trial-pooled; fraction of trials executing at least half the expected tools) and path consistency (dominant tool-sequence % among adherent trials; per-domain macro-average), 12 domains. All adherence deltas are significant. Consistency deltas carry two-stage cluster-bootstrap CIs (resampling domains, then trials within domain; 10,000 resamples): Ministral +4.6 [ − 1.8, +11.3], DeepSeek +13.7 [+5.8, +23.0], GLM 5 +7.0 [+0.1, +17.9], Kimi +4.5 [ − 0.3, +10.0]. * = CI excludes zero.
Grounded TSR
Raw TSR
Executor
Tier
RFC
MCP
Δ [95% CI]
Δ [95% CI]
Ministral
light
61.0%
67.5%
+6.5 [+3.4, +9.6] *
+2.3 [ − 0.7, +5.4]
DeepSeek
large
78.6%
79.0%
+0.4 [ − 2.2, +3.0]
− 2.5 [ − 5.1, − 0.1] *
GLM
large
84.8%
82.0%
− 2.9 [ − 5.2, − 0.5] *
− 4.7 [ − 6.9, − 2.4] *
Kimi
frontier
85.5%
83.8%
− 1.7 [ − 4.0, +0.7]
− 3.8 [ − 6.0, − 1.5] *
Table 4: Grounded and raw TSR deltas by executor (trial-pooled, 12 domains; raw = grounded + ungrounded). Bootstrap 95% CIs, 10,000 trial-level resamples; * = CI excludes zero.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Avg MCP Δ
traffic_spoofing_detection
+39.8
patient_intake
+13.3
know_your_business
+7.6
dangerous_goods
+3.1
content_flagging
+2.5
order_fulfillment
+2.5
Appendix
Table 5: Average grounded MCP lift by domain in percentage points (simple mean of the four per-model Grounded TSR deltas, computed from unrounded per-model values). Gains concentrate where prompt-based adherence was poor ( traffic_spoofing_detection , patient_intake ); losses concentrate in branching decision-tree domains where step-scoped autonomy prevents look-ahead.
rfc2119_sop
sop_mcp
Domain
M8B
DS
GLM
Kimi
M8B
DS
GLM
Kimi
aircraft_inspection
79% (25)
79% (9)
88% (15)
86% (8)
77% (28)
79% (6)
80% (4)
87% (8)
content_flagging
88% (10)
89% (5)
97% (9)
99% (4)
89% (18)
96% (3)
99% (4)
99% (4)
customer_service
58% (28)
85% (22)
92% (38)
82% (21)
70% (60)
69% (13)
88% (19)
76% (13)
dangerous_goods
57% (18)
95% (12)
96% (6)
96% (1)
78% (15)
90% (4)
94% (1)
93% (4)
email_intent
70% (33)
86% (7)
92% (8)
91% (5)
73% (16)
85% (5)
77% (6)
91% (5)
Appendix
Table 6: Grounded TSR by domain, model, and condition; parenthesized numbers are unique execution paths among adherent trials. M8B = Ministral 3 8B, DS = DeepSeek V3.2, Kimi = Kimi K2.5.
Ungrounded
Reinvention
Executor
Ungr .
/Corr
KYB
#Paths
Cons.
Ministral
4.5%
6.8%
44%
20.0
53%
DeepSeek
3.1%
3.8%
49%
8.9
65%
GLM
2.1%
2.4%
31%
10.7
70%
Kimi
2.3%
2.6%
48%
6.6
74%
Appendix
Table 7: Prompt-based delivery (RFC condition, 12 domains). Columns 2–4: correct answers produced without executing the prescribed tools (% of trials, share of correct, worst-domain share on know_your_business ). Columns 5–6: mean distinct tool-call sequences per domain among adherent trials, and dominant-sequence consistency.
RFC
MCP
Executor
Ungr/Corr
(n)
Ungr/Corr
(n)
Ministral
44%
22/50
8%
3/39
DeepSeek
49%
36/73
0%
0/35
GLM
31%
20/65
0%
0/50
Kimi
48%
34/71
0%
0/54
Appendix
Table 8: know_your_business : ungrounded correct answers as a share of all correct answers, per model and condition (90 tasks per cell). Under prompt-based delivery, 31–49% of every model’s correct answers bypass the prescribed tools; under step-level delivery this falls to 0–8%.
t
Executor
rfc2119
sop_mcp
Δ
0.3
Ministral
61.3%
67.5%
+6.3pp
DeepSeek
79.2%
79.0%
− 0.3pp
GLM
85.0%
82.0%
− 3.0pp
Kimi
86.0%
83.8%
− 2.2pp
0.5
Ministral
61.0%
67.5%
+6.5pp
DeepSeek
78.6%
79.0%
+0.4pp
Appendix
Table 9: Grounded TSR sensitivity to adherence threshold t (trial-pooled, 12 domains). Directions are stable for Ministral, GLM 5, and Kimi; DeepSeek hovers at zero.
Executor
Cond.
Adh%
Acc ∣ Adh
GTSR
Ministral
rfc2119
75.9%
80.4%
61.0%
Ministral
sop_mcp
94.7%
71.2%
67.5%
Δ
+18.8pp
− 9.1pp
+6.5pp
DeepSeek
rfc2119
86.5%
90.8%
78.6%
DeepSeek
sop_mcp
97.2%
81.3%
79.0%
Δ
+10.7pp
− 9.6pp
+0.4pp
Appendix
Table 10: Grounded TSR decomposition (trial-pooled, 12 domains). Adh% = fraction of trials passing the adherence threshold; Acc ∣ Adh = accuracy among adherent trials.
Executor
Cond.
Tok (in/out)
Dur (s)
$/trial
Ministral
rfc2119
22.9k / 0.8k
6.5
$0.004
Ministral
sop_mcp
66.0k / 1.8k
16.8
$0.010
DeepSeek
rfc2119
34.4k / 1.2k
33.4
$0.023
DeepSeek
sop_mcp
83.7k / 3.3k
93.9
$0.058
GLM
rfc2119
31.7k / 0.8k
32.3
$0.021
GLM
sop_mcp
68.9k / 1.7k
74.5
$0.045
Appendix
Table 11: Per-trial economics (trial-pooled, 12 domains). List prices (/Min/out):Ministral38B0.15/0.15,DeepSeekV3.20.62/1.85,GLM50.60/2.20,KimiK2.50.60/$3.00 (all verified against published list prices for the region used).
Executor
Cond.
Med.
p90
p99
Max
Ministral
RFC 2119
614
1,539
2,603
4,815
SOP-MCP
1,413
3,357
5,790
9,814
DeepSeek
RFC 2119
1,043
2,111
2,805
6,037
SOP-MCP
3,148
5,066
6,030
8,144
GLM
RFC 2119
756
1,462
1,828
2,137
SOP-MCP
1,586
2,791
3,753
4,391
Appendix
Table 12: Output-token usage per trial (trial-pooled, 12 domains): total generated tokens summed across all responses of a trial. Because a trial spans multiple model responses, any single response – the unit the max_tokens cap applies to – is strictly smaller than the trial total. Median usage sits at 15–77% of even the smaller 4,096 cap.
Enterprise agents must follow long-horizon, conditional, safety-critical standard operating procedures (SOPs). We compile machine-readable SOP constraints into executable pseudo-code and run them with a program-guided (PG) stack machine that pages the active frame while an LLM performs semantic execution. A three-arm SOPBench study across six models separates representation from runtime: compiled text never significantly hurts and gains up to 16.0 points where official prose underperforms. Runtime guidance is capability-gated. Two strong models independently show positive seven-domain PG contrasts (58:19 and 75:31 discordant pairs), whereas weak models are harmed. A full-program cursor ablation (active frame first, complete program retained) recovers much of the strong-model refusal gain; selective visibility adds a smaller improvement. Paired probe and audit measurements track this divide to spontaneous state discipline rather than reconstruction ability. On Bank the three primary arms rise from 70.4 to 86.4 to 92.8, with 100% refusal correctness. Practical guidance: compile first; enable active-frame paging only after a model-level discipline check.
Chenglin Yu, Li Yin, Ying Yu +4
Department of Industrial and Systems Engineering, The Hong Kong Polytechnic University · Department of Data and Systems Engineering, The University of Hong Kong · College of Economics and Management, Zhejiang Normal University +3
Large language models (LLMs) often achieve strong performance on reasoning benchmarks, but final-answer accuracy alone does not show whether they faithfully execute the procedure specified in a prompt. We introduce a controlled diagnostic benchmark for procedural execution, where models are given a step-wise arithmetic procedure and two numeric inputs, and must return the final computed value. Complexity is varied through procedure length and look-back dependencies over intermediate variables. Average first-answer accuracy drops from 63% on 5-step procedures to 20% on 95-step procedures. Generation-level analysis shows that failures often involve missing answers, premature answers, self-correction after an initial error and under-executed traces. These findings suggest that apparent reasoning ability can mask substantial weaknesses in faithful long-horizon procedural execution.
Agent orchestration frameworks -- LangGraph, CrewAI, Google ADK, OpenAI Agents SDK, and others -- place an external orchestrator above the LLM, tracking state and injecting routing instructions at every turn. We present a controlled comparison showing that for procedural tasks, this architecture is dominated by a simpler alternative: putting the entire procedure in the system prompt and letting the model self-orchestrate. Across three domains -- travel booking (14 nodes), Zoom technical support (14 nodes), and insurance claims processing (55 nodes) -- we evaluate 200 conversations per condition using LLM-as-judge scoring on five quality criteria. The in-context approach scores 4.53--5.00 on a 5-point scale while a LangGraph orchestrator using the same model scores 4.17--4.84. The orchestrated system fails on 24% of travel, 9% of Zoom, and 17% of insurance conversations, compared to 11.5%, 0.5%, and 5% for the in-context baseline. While external orchestration may have been necessary for earlier models, advances in frontier model capabilities have made it unnecessary for multi-turn conversations following a defined procedure.