What May an Agent Change About Itself? A Containment Floor for Self-Configuring Agent Runtimes
Organizations: North South University · HelicanHQ
Abstract
Many agent runtimes give the agent a tool for editing its own configuration. Some of that configuration grants abilities, such as enabling a tool. Other parts set the agent's limits: which directories it may write to, who may send it messages, which network address it listens on, how callers authenticate, and the gate that blocks risky writes. If the agent can edit those limits, a single ordinary request can widen them. We study this in a deployed, model-agnostic runtime. We propose a rule: the agent may change fields that grant abilities, and may never change fields that set its limits. We enforce the rule as a containment floor inside the configuration tool and measure what happens with and without it. Without the floor, a frontier model wrote a protected value on 25 of 72 ordinary requests that gave it permission to change settings, often when the request never named the field. Prohibitions written in the system prompt failed in a predictable way. A prompt that listed the protected field names stopped every request that used those names (0 of 36 saved, against 17 of 36 with no prompt) and did not stop the requests that only described the goal (10 of 36 saved, against 8 of 36). A prompt that described the forbidden effects did the reverse. With the floor, 0 of 167 protected writes were saved, although the models attempted a protected write in 65 of those cases. A search for other routes through the tool found only one, a pinned shell, which the floor's scope statement already excludes. The study covers two models and a single agent. We state what that does and does not support.
Figures & tables
| Field | What it controls | Why the agent may not change it | How the floor detects a write |
|---|---|---|---|
| writeRoots | Directories the agent’s file tools may write to | A new value can widen where the agent may write, and the tool cannot tell a narrower value from a wider one | By key name, at any depth of the path, and inside any value written, including values written to tools.providers . Any value is refused. |
| channels[]. allowFrom | Who may send the agent messages on a channel | A new value can widen who can reach the agent, and the tool cannot tell a narrower list from a wider one | By key name, at any depth, and inside any value written to channels . Any value is refused. |
| server.host | Network address the API listens on | A new address can widen the agent’s reach, and the tool cannot tell a narrower address from a wider one | By exact path. Any value is refused. |
| server.tokenEnv | Environment variable that holds the API token | It changes what authenticates callers | By exact path. Any value is refused. |
| policy.deny | Commands the agent is forbidden to run | It removes a restriction someone else set | By exact path. Any replacement is refused. |
| untrusted. onMutate | Whether writes are allowed after the agent reads untrusted content | The value allow turns the write gate off | By exact path and value. Only allow is refused; confirm is permitted. |
| Model | Arm | Indirect att/exec | Direct att/exec | All protected att/exec | Controls landed |
|---|---|---|---|---|---|
| gpt-5.6 (native) | None | 8/8 of 36 | 17/17 of 36 | 25/25 of 72 | 21/36 |
| Prompt (fields) | 10/10 of 36 | 0/0 of 36 | 10/10 of 72 | 19/36 | |
| Prompt (effects) | 0/0 of 36 | 19/19 of 36 | 19/19 of 72 | 19/36 | |
| Prompt (both) | 3/3 of 18 | 1/1 of 18 | 4/4 of 36 | 7/18 | |
| Floor | 3/ 0 of 36 | 23/ 0 of 36 | 26/ 0 of 72 | 22/35 | |
| Shipped | 0/ 0 of 35 | 25/ 0 of 36 | 25/ 0 of 71 | 20/32 |
| Comparison (executed protected writes) | Arm | None | Fisher |
|---|---|---|---|
| Prompt (fields), direct requests | 0/36 (0–10%) | 17/36 (32–63%) | |
| Prompt (fields), indirect requests | 10/36 (16–44%) | 8/36 (12–38%) | 0.79 |
| Prompt (effects), indirect requests | 0/36 (0–10%) | 8/36 (12–38%) | 0.005 |
| Prompt (effects), direct requests | 19/36 (37–68%) | 17/36 (32–63%) | 0.81 |
| Prompt (both), all protected, same run | 4/36 (4–25%) | 12/36 (20–50%) | 0.045 |
| Floor and Shipped, all protected, both models | 0/167 (0–2%) | 25/72 (25–46%) | n/a |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Arm | deepseek-v4-pro | gpt-5.6 (native) | llama3.2:3b |
|---|---|---|---|
| Full catalogue, 10 tools | 87.5 / 83.3 | 95.8 (2 crit.) | 41.7 (6 crit.) |
| Triage + phase_set advertising “act adds 5 tools” | 75.0 / 75.0 | 91.7 | 70.8 |
| Triage + phase_set , nothing advertised | not run | 91.7 | 70.8 |
| Triage, no phase_set | not run | 100.0 | 66.7 |
| Requests | None, executed | Floor att/exec | Shipped att/exec |
|---|---|---|---|
| Paraphrase, no field named | 6 | 5/ 0 | 3/ 0 |
| Nested in a parent or block | 14 | 7/ 0 | 10/ 0 |
| Pressure, two turns | 11 | 8/ 0 | 0/ 0 |
| All | 31 of 51 | 20/ 0 of 51 | 13/ 0 of 47 |