Multi-agent LLM systems negotiating with a stateful counterpart waste model calls in three ways: polite loops that never meet the counterpart's hidden acceptance condition, malformed outputs that trigger retries, and compliance deadlocks in which the counterpart demands something the agent must refuse. We study a three-part control stack - a 5-Pillar runtime constitution, a 4-tier swarm (Director, three-agent majority vote, Monitor, schema hard gate) and Cognitive Annealing (deterministic deadlock detection, atomic purge of the agent-side context, a canonical recovery message) - against a released adversarial Gatekeeper whose acceptance rules are fixed regular expressions and whose LLM only renders reply text. The testbed has a known solution: it measures whether the stack executes a constitution-aligned strategy against swarm drift and recovers from deadlock, not whether it discovers anything. In five runs per configuration (30 runs; Gemini 2.5 Pro agents, Claude Haiku 4.5 Gatekeeper) we find: (i) the constitution and Director make an acceptable framing possible but not reliable - 0/5 baseline unlocks versus 1/5 and 2/5 with the constitution; when the swarm unlocks it does so in one turn with 7-8 calls and about 15k tokens (67-73% below baseline); when it does not, it costs 17-38% more; (ii) the Monitor and hard gate do not reduce unlocks and leave an audit trail; (iii) under a honeytrap-to-compliance deadlock, LLM-only steering escapes 0 of 5 times while atomic purge plus a canonical strike escapes 5 of 5 (Fisher p=0.008) at the same call budget, with zero calls for the strike. LLM-written strikes failed the deterministic pre-flight 5 of 5 times although an LLM Monitor had approved four. Pre-registered hypotheses on average call and token reduction were not supported. Cost is bounded in every arm by deterministic stop rules; the stack adds recovery at no extra model cost.
Figures & tables
Figure 1: The control stack and the testbed. Left: the 5-Pillar constitution and the four tiers that turn it into one outbound message per turn, with the arms in which each tier is off. Right: the Gatekeeper’s deterministic state machine and acceptance rules; its LLM renders reply text only. Bottom: Cognitive Annealing, which runs once per conversation after the deadlock detector fires. Call counts per turn are structural.
Gatekeeper status (frustration, repetition, forced honeytrap at message 3)
state machine
—
Gatekeeper reply text
canned text on extraction failure
Claude Haiku 4.5 renders the reply for the chosen status
Director strategy, agent drafts, ballots, Monitor verdict
—
Gemini 2.5 Pro
Hard gate (Tier 3)
Pydantic schema; “?” required
rewrite call only on failure
Deadlock detector, atomic purge, canonical strike, send guard, halt
all deterministic
—
Table 2
Arm
Constitution + Director
Monitor + Tier 3
Scenario
Purge
Strike
Calls / turn
P1’
off
off
Phase 2
—
—
6
P2A’
on
off
Phase 2
—
—
7
P2B’
on
on
Phase 2
—
—
8
P3B’
on
on
Phase 3
off
off
8
P3A’-LLM
on
on
Phase 3
on
LLM candidate, send guard, canonical fallback
8 (+2 for the strike)
P3A’-CANON
on
on
Phase 3
on
canonical string
8 (+0 for the strike)
Table 3
Figure 2: Per-run model calls, swarm turns and total tokens by arm. Each dot is one run (filled: ended in MEETING_UNLOCK_APPROVED; open: ended without unlock); the bar is the median of five.
P1’
P2A’
P2B’
Unlock (of 5) [Wilson 95%]
0 [0–43%]
1 [4–62%]
2 [12–77%]
Unlock incl. interrupted attempts (Section 5.7)
0 of 6
1 of 5
2 of 7
Swarm turns
4 [4–4]
4 [1–4]
4 [1–4]
Model calls
24 [24–24]
28 [7–28]
32 [8–33]
Total tokens
55,290 [54,710–58,505]
60,613 [14,864–67,758]
71,459 [14,625–74,181]
Prompt tokens (median)
12,407
20,878
21,803
Table 5
Figure 3: Gatekeeper status after each message for all 30 runs (Gatekeeper-side numbering). S = SOFT_REJECT, F = FRAMING_LOCK, X = ACTIVE_EXPLOITATION (honeytrap), P = PROGRESSIVE_ESCALATION, U = MEETING_UNLOCK_APPROVED. A diamond marks the atomic purge, after which the strike is the next message; the right column gives model calls per run.
P3B’ (no recovery)
P3A’-LLM
P3A’-CANON
Unlock (of 5) [Wilson 95%]
0 [0–43%]
5 [57–100%]
5 [57–100%]
Deadlock detected (turn)
5 of 5 (turn 3)
5 of 5 (turn 3)
5 of 5 (turn 3; one at turn 2)
Gatekeeper phase at unlock
—
EXCEPTION_UNLOCK (5)
EXCEPTION_UNLOCK (5)
Strike actually sent
—
canonical, by fallback (5); LLM candidate accepted 0
canonical (5)
Monitor LLM verdict on the LLM candidate
—
APPROVE 4, REVISE 1 (overridden by the lexicon check in all 5)
—
Calls spent on the strike
—
2 per run, both wasted
0
Table 7
Arm
Monitor reviews
REVISE
Violations
Z-axis fail
Tier-3 pass
Tier-3 fail
Truncated
P1’
0
0
0
0
0
0
2
P2A’
0
0
0
0
0
0
1
P2B’
14
4
6
2
14
1
4
P3B’
15
0
0
0
15
2
5
P3A’-LLM
20
6
6
0
20
6
3
P3A’-CANON
14
0
0
0
14
2
4
Table 8
ID
Hypothesis
Observation
Test
Verdict
H-A
P2A’ unlocks in fewer turns than P1’
turns 4,4,4,4,4 vs 4,4,4,4,1; unlock 0 of 5 vs 1 of 5
MWU one-sided p=0.35 ; Fisher p=1.0
not supported
H-B
≥40% fewer calls
medians 24 to 28 ( −17% ; bootstrap 95% −17% to +71% ); means 24.0 to 23.8
MWU one-sided p=0.95
not supported
H-C
≥50% fewer tokens
medians 55,290 to 60,613 ( −10% ); means 55.7k to 53.0k
MWU one-sided p=0.92
not supported
H-D
Monitor + Tier 3 do not lower unlock
1 of 5 to 2 of 5 (2 of 7 with interrupted attempts)
Fisher p=1.0
consistent; low power
H-E
P3B’ never unlocks; P3A’-CANON does
0 of 5 vs 5 of 5 at equal call budget
Fisher p=0.0079
supported
H-F
P3A’-CANON unlocks more than P3A’-LLM
5 of 5 vs 5 of 5; LLM candidate accepted 0 of 5; +2 calls per run
not testable on unlock; supported on acceptance and cost
Table 9
#
Arm
Log
Start (UTC)
Outcome
Turns
Calls
Prompt
Cand.
Think.
Total
Traj.
Notes
1
P1’
P1 r1
10-05 11:09
LOOP
4
24
12,428
890
41,424
54,742
SSXP
pre-PR#2
2
P1’
P1 r3
10-07 01:38
LOOP
4
24
12,161
923
42,206
55,290
SSXP
trunc×1
3
P1’
P1 r4
10-07 05:42
LOOP
4
24
12,407
910
45,188
58,505
SSXP
4
P1’
P1 r5
10-08 02:10
LOOP
4
24
11,831
904
41,975
54,710
SSXP
5
P1’
P1 r6 retry1
10-08 05:17
LOOP
4
24
12,605
913
41,830
55,348
SSXP
trunc×1, 429 retry×1
6
P2A’
P2A r1
10-06 05:20
LOOP
4
28
20,905
1,592
38,116
60,613
SSXP
pre-PR#2
Table 10
t (s)
Turn
Event
Detail
Calls
0
STUDY_START
max_turns=8
0
0
1
TURN_START
turn 1
0
12
1
DIRECTOR_STRATEGY
strategy: Acknowledge their strong in-house team, then pivot to a hypothesis-driven question about their analyst workloa…
1
95
CONSENSUS_RESULT
winning proposal 2 (swarm calls 6)
7
103
1
MONITOR_REVIEW
APPROVE; violations=0
8
103
1
TIER3_HARD_GATE_PASS
attempt 1
8
Table 11
t (s)
Turn
Event
Detail
Calls
0
STUDY_START
max_turns=8
0
0
1
TURN_START
turn 1
0
18
1
DIRECTOR_STRATEGY
strategy: Acknowledge their stated strength in having a strong in-house team. Reframe the interaction away from a ’pitch…
1
134
CONSENSUS_RESULT
winning proposal 1 (swarm calls 6)
7
142
1
MONITOR_REVIEW
APPROVE; violations=0
8
142
1
TIER3_HARD_GATE_PASS
attempt 1
8
Table 12
t (s)
Turn
Event
Detail
Calls
0
STUDY_START
max_turns=8
0
0
1
TURN_START
turn 1
0
14
1
DIRECTOR_STRATEGY
strategy: Acknowledge their statement about having a strong in-house team, then immediately pivot to the core hypothesis…
1
104
CONSENSUS_RESULT
winning proposal 1 (swarm calls 6)
7
113
1
MONITOR_REVIEW
APPROVE; violations=0
8
113
1
TIER3_HARD_GATE_PASS
attempt 1
8
Table 13
Planned (re-run design v2, 5 Oct 2026)
Executed
Note
F1: remove “Do not pivot to orthogonal overflow-research framing.” from the baseline agent’s system note; keep only the statement that it has no constitution, Director or Monitor
as planned (commit ce6d198 )
the note still instructs “mainstream sales tactics (quality, price, meetings)”, and the product-knowledge block still names overflow research (Section 4.6)
F2: remove the “Strategic unlock (Gatekeeper mock)” paragraph from the constitution in both the Phase 2 and Phase 3 agents; Pillars 1–5 unchanged
as planned
Pillar 2 still contains “overflow / supplemental research capacity” by design
F3: replace the Director prompt’s example required_phrases / avoid_phrases with placeholders; replace “Steer toward analyst pain + overflow research framing” with “Steer toward the Constitution’s Pillar 1 and Pillar 2 framing”
as planned
the Director still produced acceptance terms on its own in some turns (e.g. “overflow capacity”)
F4: strip prompt echoes from Gatekeeper replies; on extraction failure fall back to the canned reply and record it
as planned; fired twice (Section 5.7)
written for the May Llama endpoint, kept as a safeguard with Claude
F5: record usage_metadata (prompt, candidates, total tokens) and model_version on every agent-side call; add totals to STUDY_COMPLETE
as planned; 724 call records, totals reconciled 30 of 30
total includes thinking tokens
F6: record fallback_pitch_used whenever a fallback text is sent
We propose a personal-LLM exchange (LLM-X), a scalable negotiation-oriented environment that enables direct, structured communication across populations of personal agents (LLMs), each representing an individual user. Unlike existing tool-centric protocols that focus on agent-API interaction, LLM-X introduces a message bus and routing substrate for LLM-to-LLM coordination with guarantees around schema validity and policy enforcement. We contribute: (1) an architecture for LLM-X comprising federated gateways, topic-based routing, and policy enforcement; (2) a typed message protocol supporting capability negotiation and contract-net-style coordination; and (3) the first empirical evaluation of LLM-based multi-agent negotiation at scale. Experiments span 5, 9, and 12 agents, under distinct negotiation policies (Low, Medium, High), and across both short-run (minutes) and long-run (2h, 12h) load conditions. Results highlight clear policy-performance trade-offs: stricter policies improve robustness and fairness but increase latencies and message volume. Extended runs confirm that LLM-X remains stable under sustained load, with bounded latency drift.
Giuliano Lorenzoni, Paulo Alencar, Donald Cowan
University of Waterloo · University of Waterloo Waterloo, ON, Canada
Multi-agent LLM debate improves factuality and reasoning, but most recipes pick a fixed round count, over-spending on easy items and under-spending on hard ones. We adapt Wald's Sequential Probability Ratio Test (SPRT) as a plug-in compute governor for LLM debates. After each round, an LLM judge emits a [0,1] consensus score on the latest agent positions; a Wald monitor accumulates the log-likelihood ratio of "useful convergence" vs "not yet useful" under a Beta likelihood family, and stops when either boundary is crossed or returns a capped best-effort outcome at R_max. Under i.i.d. assumptions the rule inherits SPRT type-I/type-II error guarantees; in deployment the calibration itself is the more important object, since it estimates whether the judge score actually separates useful from unhelpful convergence in a given domain. We evaluate two tracks: (i) a Monte-Carlo study under calibrated Beta models characterising working curves, error rates, capping behaviour, and sensitivity; and (ii) a real-LLM evaluation on 200 attempted MMLU and 200 attempted GSM8K items with three heterogeneous agents (gpt-5, claude-opus-4-6, gemini-2.5-pro) and a claude-opus-4-6 judge, using disjoint 40-item calibration subsets. On GSM8K the rule stops in 1.01 average rounds (4.06 LLM calls) at 97.0% accuracy vs 99.0% for fixed-5 debate at 15 calls: a 3.7x call reduction at -2pp accuracy. On MMLU the calibrated KL collapses to about 0 and the rule caps on 99.5% of items at 2.1x cost. The takeaway is not that SPRT makes debate more accurate, but that a classical sequential test serves as a cheap compute-control and failure-detection layer for multi-agent LLM systems.
Multi-agent systems built on large language models (LLMs) are difficult to reason about. Coordination errors such as deadlocks or type-mismatched messages are often hard to detect through testing. We introduce a domain-specific language for specifying agent coordination based on message sequence charts (MSCs). The language separates message-passing structure from LLM calls, tool calls, and human control points, whose outcomes remain unpredictable. We define the syntax and semantics of the language and present a syntax-directed projection that generates deadlock-free local agent programs from global coordination specifications. We illustrate the approach with a diagnosis consensus protocol and show how coordination properties can be established independently of LLM nondeterminism. We also describe a runtime planning extension in which an LLM dynamically generates a coordination workflow for which the same structural guarantees apply. An open-source Python implementation of our framework is available as ZipperGen.
Benedikt Bollig, Matthias Függer, Thomas Nowak
Universit´e Paris-Saclay, CNRS, ENS Paris-Saclay, LMF, Gif-sur-Yvette, France · Institut Universitaire de France, Paris, France