When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents
Organizations: University of Illinois Urbana-Champaign · Microsoft Research AI Frontiers
Abstract
LLM agents increasingly solve long-horizon tasks by autonomously interacting with their environment. In doing so, their strategies rely on assumptions about that environment: which resources and tools exist, where they are located, and how they behave. When these assumptions no longer hold, reliable agents must detect the change and adapt while pursuing the same goal. We study this adaptation capability through environmental novelty: a change that keeps the task objective fixed while invalidating an assumption underlying an otherwise successful trajectory. We introduce AGNI, an automated pipeline that extracts trajectory-relevant assumptions, injects targeted environmental changes, and validates that the resulting novel tasks remain solvable. Across three terminal benchmarks, AGNI produces diverse novelties spanning resources, interfaces, constraints, and execution semantics. Evaluating multiple LLM agents reveals a substantial adaptation gap between base and novel tasks. Trajectory analysis suggests that agents often encounter evidence of the change but fail to diagnose its cause and revise their strategy. Finally, post-training for environmental novelty improves adaptation to held-out novel tasks while also improving performance on base tasks. Our results highlight a gap between task competence and adaptive capability and motivate environmental variation as a core dimension of agent training and evaluation.
Figures & tables
| ET-eval | TB-Lite | TB-2 | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Base | Novel | Base | Novel | Base | Novel | |||
| GPT-5.5 | 97.0 | 74.5 | 22.5 | 93.2 | 71.7 | 21.5 | 88.8 | 79.7 | 9.1 |
| GPT-5.4-mini | 94.5 | 61.9 | 32.6 | 65.0 | 39.5 | 25.5 | 75.9 | 46.6 | 29.3 |
| Kimi-2.6 | 92.7 | 74.7 | 18.0 | 83.6 | 58.3 | 25.3 | 73.8 | 61.4 | 12.4 |
| DeepSeek-V4-Flash | 89.4 | 74.2 | 15.1 | 70.2 | 49.3 | 20.8 | 66.0 | 54.0 | 12.0 |
| GPT-OSS-120B | 78.6 | 29.4 | 49.2 | 52.7 | 17.6 | 35.1 | 28.0 | 13.3 | 14.7 |
| Model | Effort | Base | Novel | |
|---|---|---|---|---|
| GPT-5.4-mini | Low | 85.4 | 41.5 | 44.0 |
| Medium | 91.4 | 53.8 | 37.6 | |
| High | 94.5 | 61.9 | 32.6 | |
| GPT-OSS-120B | Low | 75.0 | 22.2 | 52.8 |
| Medium | 78.6 | 29.4 | 49.2 | |
| High | 84.2 | 35.3 | 48.9 |
| Method | Base | Novel | |
|---|---|---|---|
| GRPO on | 59.82 | 29.46 | 30.36 |
| + Meta Prompt | 61.25 | 32.59 | 28.66 |
| + Meta Prompt Training | 64.29 | 37.86 | 26.43 |
| + Recovery & Verify | 69.55 | 38.00 | 31.55 |
| + Diagnosed Exploration | 53.93 | 25.14 | 28.79 |
| + Exploration & Completion | 66.96 | 38.80 | 28.16 |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Role | Symbol | Purpose |
|---|---|---|
| Solver | Generates interaction trajectories in the base and novel environments. | |
| Strategy analyzer | Summarizes observed task-solving behavior and extracts trajectory-grounded environmental assumptions. | |
| Novelty generator | Proposes environmental changes that target extracted assumptions while preserving the task objective. | |
| Implementer | Instantiates a novelty specification as a modified environment and produces an adapted reference solution . | |
| Behavioral observer | Checks whether the injected novelty produces the expected change in behavior between the base and novel environments. | |
| Trajectory analyzer | Analyze Trajectories | Analyzes novel-task trajectories to characterize adaptation and identify unintended shortcuts or reward hacking. |
| Setting | Value |
|---|---|
| Model | GPT-5.5 (2026-04-24), all roles. |
| Reasoning effort | High for the solver; medium for other roles. |
| Sampling | Temperature and top- unset (API defaults); no fixed seed. |
| Output-token limits | 8,192 per solver turn; 32,000 per non-solver call. |
| Solver harness | Terminus-2 (Harbor 0.6.6); at most 32 turns. |
| Construction rollouts | per base and novel task, across all three sources. |
| Axis | Strata |
|---|---|
| Base-task difficulty | Easy, medium, hard |
| Novelty-induced difficulty | Low, medium, high |
| Occurrence plausibility | Low, medium, high |
| Realization fidelity | Low, medium, high |
| Mechanism family | Categorical labels from the frozen taxonomy |
| Adaptive behavior | Non-adaptive behavior | |
|---|---|---|
| Agent | GPT-5.4-mini, run_02 . | GPT-5.4-mini, run_07 . |
| Explanation | Reads the launch log and identifies nohup as the obstacle. | Leaves the cause unresolved and continues to rely on nohup . |
| Action | Relaunches without nohup , using python3 -m mlflow server … & . | Retries launch commands that still use nohup . |
| Check | Checks the server’s health endpoint and confirms that it responds. | Launch attempts continue to fail; no working server is established. |
| Result | Trains and registers the model; all three tests pass. | Does not complete the task; all three tests fail. |
| Adaptive behavior | Non-adaptive behavior | |
|---|---|---|
| Agent | GPT-5.4-mini, run_00 . | GPT-5.5, run_02 . |
| Explanation | Identifies the changed encoding as the reason direct searches miss the text. | Concludes that the logs contain none of the target words. |
| Action | Uses Python to decode the logs as UTF-16 before counting matching lines. | Repeats grep searches without accounting for the encoding. |
| Check | Confirms the UTF-16 encoding marker in the file bytes. | Accepts repeated zeros as evidence that the target words are absent. |
| Result | Reports counts of 4, 3, and 8; all three tests pass. | Reports zero counts; the grader rejects the summary. |
| Adaptive behavior | |
|---|---|
| Agent | GPT-5.4-mini, run_00 . |
| Explanation | Suspects an incorrect program path in the first line or Windows line endings; does not directly inspect the extra character. |
| Action | Reads the script and runs bash /app/generate_zk_protocols.sh , avoiding the broken first line. |
| Check | Confirms that the script exists and has execution permission before changing how it is run. |
| Result | Runs the generator and writes both reports; all 15 tests pass. |
| Adaptive behavior | Non-adaptive behavior | |
|---|---|---|
| Agent | GPT-OSS-120B, run_00 . | GPT-OSS-120B, run_01 . |
| Explanation | Finds that a configuration setting prevents pip from searching for packages. | Attributes the remaining failure to Python-version compatibility without confirming that explanation. |
| Action | Inspects pip configuration and removes no-index from /etc/pip.conf . | Restores pip but leaves its global configuration unchecked. |
| Check | Successfully installs requests after removing the setting. | Confirms pip’s version but accepts No matching distribution found when testing installation. |
| Result | Restores package installation; both tests pass. | Leaves installation broken and declares completion. Verifier setup also fails, so the task tests do not run. |
| Adaptive behavior | Non-adaptive behavior | |
|---|---|---|
| Agent | GPT-5.4-mini (high), run_01 . | GPT-5.4-mini (high), run_00 . |
| Explanation | Suspects that an inherited TAR_OPTIONS setting changes the stored names. | Recognizes the wrong names but does not identify the setting that changes them. |
| Action | Unsets TAR_OPTIONS when recreating and extracting the archive. | Repeats the archive command without changing the setting. |
| Check | Lists the rebuilt archive and verifies the names app.log and error.log . | Sees the raw/ prefix again but stops after extraction fails. |
| Result | Extracts the log and completes the outputs; all four tests pass. | Leaves the archive incorrect and the extracted log missing; three tests fail. |
| Adaptive behavior | Non-adaptive behavior | |
|---|---|---|
| Agent | GPT-OSS-120B, run_00 . | GPT-OSS-120B, run_02 . |
| Explanation | Identifies that the script treated a repeated header as a release. | Blames a blank line at the end of the CSV for creating an extra file. |
| Action | Reads the CSV, removes the generated files, and rebuilds the outputs while skipping header lines. | Does not read the CSV. Attempts to delete release-.json , although the extra file is named release-version.json . |
| Check | Lists the regenerated files and confirms that exactly three remain. | Sees that the count is still four but declares the task complete. |
| Result | Completes the task; all seven tests pass. | Leaves the extra file and log entry; two tests fail and five pass. |
| Adaptive behavior | Non-adaptive behavior | |
|---|---|---|
| Agent | GPT-5.5, run_01 . | GPT-5.5, run_03 . |
| Explanation | Reads the startup error and identifies an environment setting that requests a missing database script. | Treats a successful build and the expected JAR path as sufficient to declare completion. |
| Action | Sets the database connection explicitly in the application and rebuilds the JAR. | Configures the database through the properties file, which the environment setting overrides. |
| Check | Launches the rebuilt JAR, confirms startup, and tests login. | Checks that the JAR exists, but never launches it or tests an endpoint. |
| Result | Runs the service successfully; all ten tests pass. | The grader cannot start the service; all ten tests report setup errors. |
| Adaptive behavior | Non-adaptive behavior | |
|---|---|---|
| Agent | Kimi-2.6, run_03 . | Kimi-2.6, run_04 . |
| Explanation | Recognizes that Git’s saved decision has left the requested biography update missing. | Recognizes that Git kept the old biography, but treats the conflict as resolved. |
| Action | Retrieves the updated biography from the project history and saves it. | Saves the result with the old biography still in place. |
| Check | Reads the biography and checks whether it includes the requested Stanford update. | Reads the old biography but accepts Git’s resolved status as sufficient. |
| Result | Restores the updated biography and page layout; both tests pass. | Restores the layout but leaves the biography unchanged; one test fails and one passes. |
| Score | Occurrence plausibility | Realization fidelity |
|---|---|---|
| 1 | Implausible or specific to the evaluator | Unfaithful or incoherent implementation |
| 2 | Technically possible but contrived | Crude proxy with major artificial artifacts |
| 3 | Plausible in some deployments | Controlled abstraction preserving the central effect |
| 4 | Recognizable real deployment or configuration change | Faithful simulation with minor simplifications |
| 5 | Documented, observed, or widely recognized failure mode supported by the artifacts | Direct or near-exact reproduction of the real mechanism |
| Occurrence plausibility | Realization fidelity | ||||||
|---|---|---|---|---|---|---|---|
| Dataset | Pairs | Counts (1–5) | Mean | Counts (1–5) | Mean | ||
| ET-eval | 140 | 0, 2, 43, 86, 9 | 3.73 | 68% | 0, 0, 20, 35, 85 | 4.46 | 86% |
| TB-Lite | 118 | 0, 6, 25, 76, 11 | 3.78 | 74% | 0, 1, 11, 41, 65 | 4.44 | 90% |
| TB-2 | 87 | 0, 2, 18, 60, 7 | 3.83 | 77% | 0, 0, 12, 29, 46 | 4.39 | 86% |
| All | 345 | 0, 10, 86, 222, 27 | 3.77 | 72% | 0, 1, 43, 105, 196 | 4.44 | 87% |
| Model | Provider | Model / snapshot | Main effort | Other efforts |
|---|---|---|---|---|
| GPT-5.5 | OpenAI | gpt-5.5 (2026-04-24) | High | – |
| GPT-5.4-mini | OpenAI | gpt-5.4-mini (2026-03-17) | High | Medium, Low |
| GPT-OSS-120B | OpenAI | gpt-oss-120B | Default | High, Low |
| Kimi-2.6 | Moonshot | Kimi-K2.6 (2026-04-20) | Default | High, Low, None |
| DeepSeek-V4-Flash | DeepSeek | DeepSeek-V4-Flash-0731 (2026-07-31) | None | – |
| Grok-4.20 | xAI | grok-4.20-beta-0309 | Non-reasoning | – |
| Setting | Value |
|---|---|
| Agent harness | Terminus-2, at most 32 interaction turns. |
| Rollouts | per model–task pair, unless otherwise specified. |
| Completion budget | 8,192 tokens per model response. |
| Sampling | Temperature and top- unset (deployment defaults); no explicit seed. |
| Verifier timeout | the configured verifier timeout. |
| ET-eval | TB-Lite | TB-2 | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Interact@1 | Interact@1 | Interact@1 | ||||||||||
| Model | Base | Novel | Bypass@1 | Fail w/o Int. | Base | Novel | Bypass@1 | Fail w/o Int. | Base | Novel | Bypass@1 | Fail w/o Int. |
| GPT-5.5 | 99.1 | 99.0 | 0.7 | 1.2 | 96.9 | 96.3 | 1.7 | 7.2 | 96.6 | 98.3 | 1.4 | 1.4 |
| GPT-5.4-mini | 95.3 | 95.9 | 2.7 | 3.8 | 92.7 | 94.9 | 1.5 | 8.6 | 91.2 | 95.0 | 0.7 | 8.1 |
| Kimi-2.6 | 97.3 | 98.4 | 1.6 | 0.0 | 94.2 | 97.8 | 0.8 | 3.3 | 92.6 | 96.1 | 2.3 | 4.2 |
| DeepSeek-V4-Flash | 96.1 | 97.8 | 1.3 | 3.4 | 95.4 | 96.6 | 0.7 | 5.4 | 94.6 | 95.5 | 0.8 | 4.2 |
| Retry prevalence | Retry persistence | |
|---|---|---|
| Model | Rollouts with Retry (%) | (%) |
| GPT-5.5 | 9.8 | 9 |
| GPT-5.4-mini | 15.2 | 50 |
| Kimi-2.6 | 18.1 | 6 |
| DeepSeek-V4-Flash | 17.4 | 10 |
| GPT-OSS-120B | 27.8 | 40 |
| Stage | Definition |
|---|---|
| Interacted | An executed command touches an affected tool, path, or file. Detected deterministically; no visible effect is required. |
| Observed | Terminal output exposes novelty-related evidence, such as an error, unexpected result, or changed file state. Recognition is not required. |
| Recognized | The agent explicitly identifies the observation as unexpected or inconsistent with its assumptions, without necessarily identifying the cause. |
| Diagnosed | The agent identifies the changed condition sufficiently to explain the problem and guide a response. Restating a symptom is insufficient. |
| Revised | A changed strategy successfully removes or bypasses the obstacle while preserving the task. Failed attempts and exploits do not count. |
| Solved | The agent completes the original task under novelty using a task-preserving strategy. |
| Model | Diag. Rev. | Diag. Retry | Premature Done | Pass Diag. | Pass Rev.+Diag. | Pass Rev. w/o Diag. |
|---|---|---|---|---|---|---|
| GPT-5.5 | 71 | 5 | 14 | 96 | 96 | 68 |
| Kimi-2.6 | 41 | 4 | 11 | 86 | 87 | 50 |
| DeepSeek-V4-Flash | 47 | 5 | 12 | 86 | 86 | 70 |
| GPT-5.4-mini | 57 | 11 | 20 | 87 | 86 | 59 |
| GPT-OSS-120B | 30 | 10 | 21 | 65 | 64 | 22 |
| Grok-4.20 | 31 | 14 | 16 | 61 | 64 | 20 |
| Setting | Value used in our experiments |
|---|---|
| Optimizer | AdamW; , |
| Weight decay | |
| Learning rate | , constant; no warmup |
| Gradient clipping | |
| Batch size | |
| Rollouts per prompt |
| ET-eval | TB-Lite | |||
|---|---|---|---|---|
| Qwen3-14B GRPO | Base | Novel | Base | Novel |
| Clean | 65.3 | 24.6 | 17.2 | 2.9 |
| Novel-paired | 67.1 | 35.5 | 18.9 | 5.5 |
| Prompt | Purpose | Usage in Algorithm 1 |
|---|---|---|
| Defines novelty for reuse across prompts | Lines 3, 4, 6, 11, 16: shared model context | |
| Characterizes the base strategy and assumptions | Line 3: | |
| Identifies applicable novelty dimensions | Line 4: | |
| Synthesizes novelty specifications | Line 6: | |
| Instantiates and adapted | Line 11: | |
| Repairs a failed task instantiation | Line 11: retry (implicit) |