Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents
Organizations: Microsoft
Abstract
When a tool-using agent's write times out or returns a server error, the action may already have taken effect. Retrying blindly duplicates it -- a second charge, a second announcement, a second deployment -- while giving up skips required work. We ask where exactly-once behaviour should be enforced: in the model, in the agent harness, or in the tool contract. We introduce LIMBO, a deterministic sandbox of six services with realistic contracts (optional idempotency keys, eventually consistent and missing read paths) and twelve fault modes injected at the service boundary, including late commits, redelivery and partial batches; every episode is graded against a ledger of committed effects. Across 25,930 episodes spanning nine recent models, three production agent harnesses, two contract variants and fifteen recovery conditions, the answer depends on the fault. When an immediate read-back can reveal what happened, the model decides: frontier models instructed to act exactly once almost never duplicate a write whose acknowledgement was lost (0.5%), weaker models often do, and the model explains 53% of the explained variance. When it cannot -- the request is still in flight, or the transport delivered it twice -- the same frontier models duplicate in 56% and 74% of episodes, and the contract explains 81%. We prove that no verification-only policy is exactly-once under late commits without a bound on in-flight time. Waiting works when such a bound is short and known, but with heavy-tailed in-flight delays even an hour of waiting per episode falls short of offering an idempotency key on every write, which lowers the duplicate rate from 28% to 4% because agents use keys when they exist. The harness barely matters, a guard that attaches keys transfers across harnesses unchanged, and agents reported success in 90% of the episodes in which they had duplicated an effect.
Figures & tables
| Service | Writes (idempotency) | Read-back path (consistency) |
|---|---|---|
| social | publish (key on mastodon only) | list_posts : weibo lags 180 s; none for x |
| billing | create_charge (key), refund (naturally idempotent) | list_charges (strong) |
| tickets | create , add_comment (non-idempotent), update_status (conditional) | list_recent (strong), search (lags 120 s) |
| send (non-idempotent, irreversible) | search_sent (lags 120 s) | |
| data | insert , insert_many (non-atomic), upsert (idempotent) | query (strong) |
| deploy | trigger (non-idempotent) | list_runs , get_run (strong) |
| Mode | Agent observes | Hidden outcome | Exactly-once recovery |
| timeout_pre | timeout | not executed | verify, then re-issue |
| timeout_post | timeout | executed, response lost | verify, then skip |
| timeout_late | timeout | executed 90 s later (still in flight) | same key, or wait and escalate |
| timeout_late_tail | timeout | executed 40 s–2 h later (heavy tail) | same key |
| http500_pre | HTTP 500 | not executed | verify, then re-issue |
| http500_post | HTTP 500 | executed | verify, then skip |
| Model | Duplicate rate (%) under committed faults | TS (%) when | EOS (%) | ||
|---|---|---|---|---|---|
| lost ack / 500 / partial | late commit | redelivery | not executed | all faults | |
| gpt-6-astra | 0 [0,9] | 25 [13,42] | 74 [57,86] | 100 [91,100] | 79 [66,89] |
| gpt-6-sol | 0 [0,9] | 46 [30,63] | 74 [57,86] | 99 [89,100] | 77 [63,87] |
| gpt-5.6-sol | 1 [0,10] | 65 [48,79] | 74 [57,86] | 98 [88,100] | 75 [61,85] |
| claude-opus-5.5 | 0 [0,9] | 68 [51,82] | 74 [57,86] | 100 [91,100] | 72 [58,83] |
| gemini-3.8-flash | 14 [6,29] | 59 [42,75] | 74 [57,86] | 100 [91,100] | 69 [54,80] |
| Factor | pooled (prereg.) | read-back resolves | read-back cannot |
|---|---|---|---|
| contract share (%) | 36 | 30 | 81 |
| mode share (%) | 56 | 18 | 10 |
| model share (%) | 9 | 53 | 8 |
| total pseudo- | 0.66 | 0.51 | 0.72 |
| duplicate rate (%) | 38 | 11 | 66 |
| episodes | 3,312 | 1,696 | 1,616 |
| Contract | Condition | timeout (committed) | 500 (committed) | timeout (late commit) | partial batch | redelivery |
|---|---|---|---|---|---|---|
| native | vanilla | 9 | 21 | 61 | 25 | 74 |
| native | guard | 0 | 0 | 68 | 0 | 74 |
| keys-everywhere | vanilla | 4 | 6 | 9 | 0 | 7 |
| keys-everywhere | guard | 0 | 0 | 7 | 0 | 0 |
| fixed delay (90 s) | heavy-tailed delay (40 s–2 h) | ||||||
| Condition | EOS (%) | dup. (%) | time (min) | EOS (%) | dup. (%) | time (min) | |
| vanilla | 39 | 61 | 2.9 | 34 | 66 | 2.8 | – |
| wait 0 s | 38 | 62 | 2.8 | 36 | 64 | 2.8 | 0 |
| wait 60 s | 43 | 57 | 3.2 | 42 | 58 | 2.9 | 12 |
| wait 120 s | 99 | 1 | 3.8 | – | – | – | – |
| wait 300 s | 99 | 1 | 6.2 | 67 | 33 | 6.1 | 50 |
| Model | vanilla | aware | reflect | sdk-retry | rules | vbr | guard | state oracle | outcome oracle |
|---|---|---|---|---|---|---|---|---|---|
| EOS (%) | |||||||||
| gpt-6-sol | 79 | 78 | 78 | 51 | 80 | 78 | 76 | 82 | 88 |
| claude-opus-5.5 | 77 | 77 | 77 | 51 | 77 | 77 | 77 | 82 | 88 |
| gemini-3.8-flash | 74 | 77 | 71 | 51 | 74 | 73 | 76 | 81 | 88 |
| mai-code-1.1-flash | 59 | 67 | 58 | 48 | 59 | 62 | 76 | 74 | 83 |
| Duplicate rate (%) | |||||||||
| Harness | Model | Duplicate rate (%), native | EOS (%) native | EOS (%) guard | n | ||
| committed | late | redelivery | |||||
| minimal | gpt-6-sol | 0 [0,7] | 38 [21,57] | 75 [55,88] | 77 [69,83] | 72 [63,79] | 266 |
| minimal | gpt-5.6-sol | 0 [0,7] | 67 [47,82] | 75 [55,88] | 71 [62,78] | 73 [64,80] | 266 |
| minimal | claude-opus-5.5 | 0 [0,7] | 62 [43,79] | 75 [55,88] | 73 [64,80] | 73 [64,80] | 266 |
| minimal | gemini-3.8-flash | 12 [6,24] | 58 [39,76] | 75 [55,88] | 69 [60,76] | 71 [62,78] | 266 |
| copilot | gpt-5.6-sol | 2 [0,11] | 71 [51,85] | 75 [55,88] | 70 [62,78] | 71 [62,78] | 266 |
| native contract | keys everywhere | EOS (%) | |||
| Harness | vanilla | guard | vanilla | guard | keys + guard |
| minimal | 35 | 34 | 0 | 0 | 100 |
| copilot | 37 | 36 | 0 | 0 | 100 |
| hermes | 35 | 31 | 0 | 0 | 100 |
| codex | 36 | 37 | 0 | 0 | 100 |
| Fault family | metric | default (%) | plain (%) | (pp) | McNemar | pairs |
|---|---|---|---|---|---|---|
| not executed (TS) | TS | 98 | 100 | +1.5 | 0.25 | 204 |
| read-back resolves | dup. | 12 | 22 | +9.7 | 1.5e-09 | 414 |
| late commit | dup. | 58 | 71 | +12.7 | 2.6e-06 | 204 |
| redelivery | dup. | 74 | 75 | +1.5 | 0.25 | 204 |
| no fault (EOS) | EOS | 100 | 100 | +0.0 | 1 | 72 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Tool | Kind | Description shown to the agent |
|---|---|---|
| social_publish | write | Publish a post to a social platform (mastodon, weibo, linkedin, x). Returns the new post_id. idempotency_key is honored by mastodon only; other platforms ignore it. On weibo, newly published posts can take up to 3 minutes to appear in social_list_posts. |
| social_list_posts | read | List the most recent posts on a platform (newest first). Not available for x. Weibo listings are eventually consistent (new posts may take up to 3 minutes to appear). |
| social_delete_post | write | Delete a post by id. |
| billing_create_charge | write | Charge a customer’s saved payment method. amount_cents is an integer number of cents. Returns the charge. |
| billing_list_charges | read | List a customer’s charges (newest first), including refunded ones. |
| billing_refund_charge | write | Fully refund a charge. Refunding an already-refunded charge returns 409. |
| Prediction | Result | Verdict | |
|---|---|---|---|
| H1 | Non-idempotent writes with an eventual or missing read path are duplicated in 10% of committed-fault episodes, more often than keyable or strongly verifiable writes | mixed-effects OR 6.4 (5.2–7.9); GEE OR 2.9 (1.4–6.0), | supported |
| H2 | The contract explains more variance than the model; the harness explains 10% | pooled shares: contract 36%, model 9%, harness 0% (E3); stratified: model 53% on resolvable faults, contract 81% on unresolvable ones | contract model supported only pooled; harness not supported |
| H3 | Given verification, eventual read paths yield more duplicates than strong ones | 13.4% vs 0.8% ( ), GEE | supported |
| H4 | Blind re-issue does not differ between irreversible and reversible writes (TOST, 5 pp) | difference +3.4 pp, 90% interval -1.6 to +8.5 pp | inconclusive |
| H5 | 20% of duplicate-producing episodes end “completed” with no uncertainty listed | 80% ( ); overclaim 90% (78–96%) | supported |
| H6 | The guard in the MCP server brings duplicates to 2% in every harness with 3 pp TS loss | native contract: 29% (Copilot CLI), 28% (Hermes), 29% (Codex CLI) | not supported |
| Guard vs. | EOS gain (pp) | paired episodes | Holm-adjusted |
|---|---|---|---|
| vanilla | +4.0 | 820 | 0.00011 |
| aware | +1.6 | 820 | 0.035 |
| reflect | +5.4 | 820 | 1e-07 |
| sdk-retry | +26.2 | 820 | 9.1e-59 |
| rules | +3.9 | 820 | 0.00015 |
| vbr | +3.8 | 820 | 5.9e-05 |