cs.CLOct 6, 2026

One Step at a Time: Trading LLM Autonomy for Process Predictability

Authors: Hans Schabert, Christoph Peters

Organizations: Amazon Web Services · University of the Bundeswehr Munich

Abstract

Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step. When an agent is the executor that predictability is normally lost: the prescribed procedure goes into the system prompt, and only a final answer comes back. We deliver the procedure step by step over the Model Context Protocol (MCP) instead: a server releases one step at a time, the agent executes it, and each step returns a structured step_output. This trades autonomy for predictability, and two properties then follow by construction, independent of the executor. The execution path is prescribed before the run, so the process is predictable in advance rather than reconstructed afterwards; and the completed step records form a machine-readable execution log that downstream tooling can audit and optimize step by step. Evaluating 15,475 trials across 13 SOP-Bench domains and four open-weight executors from frontier (Kimi K2.5) to lightweight (Ministral 3 8B), we find step-level delivery makes the executed process predictable and inspectable for every executor, and additionally raises accuracy when the executor is small. Across all four, process adherence rises significantly (76-95% to 95-99%) and ungrounded answers (correct outputs produced without executing the SOP) near-vanish, falling from 2.1-4.5% to 0.2-0.3% of trials (all 95% CIs exclude zero); under prompt-based delivery, 31-49% of correct answers on know_your_business bypass the SOP entirely, even for the frontier executor. Accuracy is where the executor's capability enters: the lightweight executor gains +6.5pp grounded accuracy because supplying the process externally removes a reconstruction burden it cannot carry, while capable ones trade a small raw-accuracy decrement for a predictable, auditable process.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents

    Jul 13, 2026Chenglin Yu, Li Yin, Ying Yu +4Instruction FollowingRuntime Enforcement for AI Agents

  2. When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models

    May 1, 2026Sailesh Panda, Pritam Kadasi, Abhishek Upperwal +1LLM EvaluationInstruction Following

  3. In-Context Prompting Obsoletes Agent Orchestration for Procedural Tasks

    Apr 30, 2026Simon Dennis, Michael Diamond, Rivaan Patil +2LLM PromptingLLM Agent Orchestration