Shared and structured inputs undermine collective random choice by reasoning AI agents
Organizations: Research Center for Advanced Science and Technology, The University of Tokyo 4-6-1 Komaba, Meguro-ku, Tokyo 153-8904, Japan · Department of Aeronautics and Astronautics, School of Engineering, The University of Tokyo 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan
Abstract
Random selection is widely used in resource allocation and auditing, making reliable implementation essential for AI-agent systems. Behavioural tests across six reasoning models uncovered threshold and divisibility rules used in identifier-based choices. For threshold-following GPT-6 Sol and Gemini 3.8 Flash, single-agent measurements prospectively predicted correlated participation under shared identifiers and biased participation under distinct identifiers with common timestamp bits. Changing dates, formats and identifier labels revealed when these predictions held. Explicit instructions to randomize independently reduced but did not eliminate shared-input correlation. To test implications for oversight, we asked four models to select customer requests randomly for human review. GPT-6 Sol approached the target rate while selecting predictably from identifiers; the others rarely selected requests. All four closely followed supplied random draws. These findings expose collective and audit vulnerabilities that selection rates alone miss, making input-dependent bias, correlation and predictability central targets for agent evaluation.
Figures & tables
| Scientific question | Tests | Principal evidence | Limits or predictions not met |
|---|---|---|---|
| Which identifier rules describe individual choices? | Study 1: six models, ten settings (seven with reasoning) | All seven reasoning settings depended on the identifier. Mean threshold effects over the four targets were 0.62–0.80; the remaining reasoning settings used residue or mixed rules. | All four threshold settings fell below the predicted effect of 0.4 at . Luna was mixed, rather than the predicted residue type. |
| Do individual rules predict group outcomes? | Study 2: Sol, Gemini and Luna; groups of four | For Sol and Gemini, shared or time-ordered identifiers raised loss to 3.5–6.9 times that of independent draws at . Input-feature predictions beat marginal predictions (mean log-score gains +0.78, +0.79). | Counter-identifier forecasts failed. Luna showed weaker dependence. Mixed-model groups were not tested. |
| Does creation date drive rare actions? | Study 3: Sol, Gemini and Luna; dates, targets and UUIDv7 formats | With raw UUIDv7s, participation followed the predicted date–target staircase (threshold effects 0.95, 0.41). | Hyphens reduced participation; Gemini’s rates were below the point forecasts. |
| Does the audit policy yield random selection at the requested rate? | Studies 4 and 5: Sol, Gemini, Luna and Claude | With ordinary UUIDv4s, Sol selected for review 9 of 1,008 requests (0.89%), all satisfying one of two identifier rules (one found after data collection). Gemini and Claude selected for review 0 and 2 of 1,008, respectively. | Predicted over-review of hyphenated UUIDv7s failed. Luna’s upper audit-rate bound was 1.01%, missing the below-1% criterion. |
| Do identifier labels or instructions to randomize change the effect? | Studies 6 and 7: Sol and Gemini; prompt controls and matched groups | Changing agent_id to record_id reduced the shared-minus-private loss difference by 0.078 and 0.169. Explicit randomization instructions still left differences of 0.120 and 0.057. | Shared record_id still raised Sol’s loss; Gemini largely stopped participating. The earlier joint wording and label test missed its threshold-effect criterion in both models. |
| Do audit rules persist on new requests? | Study 7: Sol; new designed and ordinary UUIDv4 requests | With new ordinary UUIDv4s, Sol selected for review 17 of 18 requests below the 1% value threshold and none of the other 990. | The last-16-bit rule failed its criterion on new requests (effect 0.06, CI to 0.17). |
| # | Role | Study | Reading | Setting | Result | Outcome |
|---|---|---|---|---|---|---|
| 1 | P | Stage 1 | (lower bound of 98.3% CI) | Sol | 0.029 [0.016, 0.043] | met |
| 2 | P | Text run 1 | Sol | lower bound 0.030 | met | |
| 3 | S | Text run 1 | Japanese rule difference positive (one-sided ; predicted ) | Sol | 0.43 (one-sided ) | met |
| 4 | S | Text run 1 | Japanese rule difference without reasoning | Sol off | met | |
| 5 | S | Text run 1 | English rule difference | Sol | met | |
| 6 | Text run 2 | (replication) | Sol | lower bound 0.018 | met |
| # | Role | Study | Reading | Setting | Result | Outcome |
|---|---|---|---|---|---|---|
| 68–69 | S | 7‡ | agent_id : shared minus private identifier loss, lower bound | Sol; Gemini 3.8 | +0.172 [0.132, 0.214]; +0.173 [0.136, 0.213] | met (2) |
| 70–71 | P | 7‡ | label interaction: shared-minus-private loss under record_id minus that under agent_id , upper bound | Sol; Gemini 3.8 | 0.078 [ 0.123, 0.034]; 0.169 [ 0.215, 0.125] | met (2) |
| 72 | S | 7‡ | record_id : shared minus private identifier loss, lower bound | Sol | +0.093 [0.062, 0.127] | met |
| 73–74 | P | 7‡ | own agent_id with shared record_id : loss minus private loss minus half the shared-identifier excess, upper bound , and own-identifier threshold effect, lower bound | Sol; Gemini 3.8 | 0.072 [ 0.096, 0.049], own-identifier effect 0.76 [0.66, 0.83]; 0.078 [ 0.099, 0.057], own-identifier effect 0.72 [0.61, 0.80] | met (2) |
| 75–76 | S | 7‡ | UUIDv7s: participation under agent_id minus under record_id , lower bound | Sol; Gemini 3.8 | 0.395 [0.328, 0.461]; 0.797 [0.746, 0.848] | met (2) |
| 77–78 | S | 7‡ | instruction to choose at random, private identifiers: threshold effect, lower bound | Sol; Gemini 3.8 | 0.67 [0.55, 0.75]; 0.38 [0.28, 0.49] | met; not met |
| Study | Data collection | Models and settings | Requests | Unanswered or invalid | Cost |
|---|---|---|---|---|---|
| Stage 1 | 24 Sep 14:00–25 Sep 01:52 | Sol, effort none and high; Japanese | 10,240 | none | 38.63 |
| Screen | 25 Sep 04:57–05:42 | Sol, Luna, GPT-5.4 mini, Claude Sonnet 5, Claude Haiku 4.5, Gemini 3.8 Flash, Gemini 3.5 Flash; reasoning on and at the lowest setting | 3,264 sent (28) | 8 invalid; Haiku 4.5 excluded at the technical stage (544 not sent) | 26.21 |
| Text run 1 | 25 Sep 08:08–08:31 | Sol high (and none); Japanese and English | 2,016 (13) | none | 9.02 |
| Text run 2 | 25 Sep 09:39–10:17 | Sol high; five languages | 2,800 (15) | none | 17.02 |
| Text run 3 | 25 Sep 11:39–11:57 | Sol high; three Japanese wordings and English | 1,280 (9) | none | 7.88 |
| 1 | 25 Sep 14:31–14:56 | 10 settings (Methods) | 3,968 (20) | 2 unanswered (HTTP 502, 503) | 29.55 |
| Loss | Unanimity | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Condition | Particip. | Observed | Input | Marginal | Observed | Input | Marginal | Gain |
| Sol (high) | no ID | 0.24 | 0.067 [0.053, 0.081] | 0.058 | 0.066 | 0.39 | 0.41 | 0.45 | 0.03 [ 0.07, 0.01] |
| private ID | 0.36 | 0.041 [0.029, 0.054] | 0.049 | 0.056 | 0.13 | 0.18 | 0.21 | +0.53 [0.23, 0.78] | |
| private number | 0.33 | 0.051 [0.036, 0.070] | 0.054 | 0.056 | 0.20 | 0.18 | 0.21 | +0.99 [0.87, 1.14] | |
| counter IDs | 0.57 | 0.075 [0.058, 0.091] | 0.369 | 0.056 | 0.00 | 0.74 | 0.21 | 1.51 [ 1.94, 1.05] | |
| shared ID | 0.36 | 0.211 [0.172, 0.250] | 0.197 | 0.191 | 0.91 | 0.85 | 0.85 | +0.64 [0.51, 0.77] | |
| Setting | Rule (study 1) | Agreement | Tokens | Shared ID | Dates | ID effect | Natural IDs | Record ID | Summaries |
|---|---|---|---|---|---|---|---|---|---|
| GPT-6 Sol, high | threshold (0.80, 0.06) | 0.92 | 126 | 0.211 (0.91) | 0.95 | 0.79 | 9 of 1,008 | 0.41 | 314 of 640 |
| GPT-6 Sol, none | no rule (0.02, 0.07) | – | 0 | – | – | – | – | – | – |
| GPT-6 Luna, high | weak or mixed (0.31, 0.23) | 0.63 | 636 | 0.092 (0.33) | 0.27 | 0.15 | 4 of 1,008 | – | not requested |
| GPT-6 Luna, none | no rule ( 0.03, 0.05) | – | 0 | – | – | – | – | – | – |
| GPT-5.4 mini, high | residue (0.02, 0.32) | 0.60 | 4,042 | – | – | – | – | – | not requested |
| Gemini 3.8 Flash, LOW | threshold (0.68, 0.05) | 0.88 | 259 | 0.192 (0.88) | 0.41 | 0.04 | 0 of 1,008 | 0.28 | 638 of 639 |
| Model | Identifier condition | Loss | Squared bias | Independent variance | Covariance component |
|---|---|---|---|---|---|
| Sol (high) | Private | ||||
| Shared | |||||
| Time-ordered UUIDv7 | |||||
| Gemini 3.8 (LOW) | Private | ||||
| Shared | |||||
| Time-ordered UUIDv7 |
| System | Identifier | Structure | Source |
|---|---|---|---|
| GitHub (MCP server) | issue and pull request numbers | integer per repository | https://github.com/github/github-mcp-server |
| Jira (Rovo MCP server) | issue keys (ABC-123) | project key and project-level counter; distinct from internal numeric ID | https://support.atlassian.com/atlassian-ai-gateway/docs/use-rovo-search-and-fetch-in-the-atlassian-remote-mcp-server/ ; https://support.atlassian.com/jira/kb/how-to-get-issue-id-from-the-jira-user-interface/ |
| Zendesk | ticket ID | sequential | https://support.zendesk.com/hc/en-us/articles/4408886154906 |
| ServiceNow | INC and CHG numbers | prefix and counter | https://www.servicenow.com/docs/r/zurich/platform-administration/c_ManagingRecordNumbering.html |
| Slack | message ts | channel-unique message ID, partly based on epoch seconds | https://docs.slack.dev/messaging/retrieving-messages/ |
| Discord, X | snowflake IDs | time, worker and sequence bits | https://docs.discord.com/developers/reference ; https://docs.x.com/fundamentals/x-ids |
| Run | Date | Purpose | Records | Sent | Cost | Reported |
|---|---|---|---|---|---|---|
| api_audit, api_audit_network | 23 Sep | interface checks | 8 | 8 | 0.00 | no |
| p1 | 23 Sep | first pilot (P1), four models | 5,280 | 5,280 | 10.63 | Note 6 |
| f0_v11_sol | 23 Sep | technical pilot for F1 | 96 | 96 | 0.26 | no |
| f1_v11_sol | 23 Sep | exploratory study F1 | 21,376 | 21,376 | 54.57 | Note 6 |
| luna_first, luna_deep | 24 Sep | exploratory Luna screens | 2,240 | 2,240 | 0.88 | Note 6 |
| delegation_stage1, _knowledge | 24 Sep | unrelated question | 204 | 174 | 5.39 | no |