cs.AIOct 1, 2026

Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay

Authors: Haotian Chen, Bowen Ye, Yuning Zhang, Jingkun Yu

Organizations: School of Cyber Science and Technology, University of Science and Technology of China · School of Education, Shanghai Jiao Tong University · SWJTU-Leeds Joint School, Southwest Jiaotong University

Abstract

Concurrent actions in large language model (LLM) agent environments require arbitration even when each proposal is individually valid. We implement a typed snapshot-settlement contract and audit three distinct properties: order sensitivity, useful progress, and replay consistency. Five settlement policies are tested in 28,800 exhaustive permutation trials and 2,160 scripted multistep episodes. Joint policies are spatially order-invariant conditional on fixed priorities, yet conservative rejection completes only 31.25% of agents in a six-agent doorway task versus 90.28% for random tickets; the paired improvement is 59.03 percentage points (95% bootstrap interval: 50.00-68.06). All policies preserve the tested spatial constraints, and priority arbitration still misses the independent small-instance optimum. A separate full-state journal audit exactly replays 156 checkpoints and rejects 1,332 constructed corruptions with a retained terminal anchor. The evidence concerns execution semantics, not human realism or long-run fairness.

Figures & tables

Explore similar work

CardsList
  1. Auditing Emergent LLM-Agent Collaboration through Cooperation-Obligation Coupling

    Jul 29, 2026Zuyuan Zhang, Hanqing Yang, Carlee Joe-Wong +1Model AuditingPassive

  2. Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures

    Sep 28, 2026Dong Xu, Zhangfan Yang, Jiantao Wu +5Task Success RateLarge Language Model Agents

  3. How Strongly Should Task State Influence an LLM Agent?

    Sep 22, 2026Chenyu Zhang, Wonbin Kweon, Jiawei HanLarge Language Model AgentsState-Tracking