From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation
Organizations: Nanyang Technological University · The Hong Kong Polytechnic University
Abstract
Realistic environment replicas are increasingly valuable for training and evaluating LLM agents, yet the original systems may be inaccessible or impractical to reproduce. We explore agentic language world modeling: rather than rebuilding an executable environment, a world model agent serves as the environment for a task agent and supports faithful and stateful simulation. We instantiate this paradigm with Trace2Env, a learning-free framework for settings where the original system is unavailable but historical interaction traces remain accessible. Trace2Env reconstructs these traces into a reusable environment worldbook containing environment schemas, grounded evidence, and induced behavioral knowledge. At runtime, the world model agent actively consults the worldbook together with persistent episodic state to infer each action's observation and lasting state effects. Across nine environments, Trace2Env improves both next-observation fidelity and long-horizon interaction consistency over conventional prompt-based LWMs. In multi-turn interaction, task agent actions generated against Trace2Env remain valid more often when replayed in the real environment, indicating that its simulated dynamics better preserve the consequences of earlier actions across successive turns. These results establish agentic language world modeling as an alternative direction for building realistic environment replicas without reconstructing the original executable system.
Figures & tables
| Method | Environment Score | Avg. | ||||||
| Terminal | SWE | Android | Web | Food | Shopping | Benefits | ||
| World-Model Backbone: gpt-5.6-sol | ||||||||
| Direct Prompting | 55.76 | 65.87 | 62.80 | 54.23 | 81.67 | 77.32 | 88.89 | 69.51 |
| Trace RAG Prompting | 58.42 | 66.50 | 66.50 | 57.12 | 85.57 | 86.07 | 96.11 | 73.76 |
| Worldbook Prompting | 59.07 | 67.82 | 65.53 | 57.70 | 86.74 | 85.89 | 97.68 | 74.35 |
| Harness only | 60.06 | 69.89 | 62.62 | 55.40 | 82.07 | 77.62 | 87.89 | 70.79 |
| World Model | Real | WM | W2R | CR |
| ALFWorld | ||||
| Direct Prompting | 93% | 97% | 3% | 0.032 |
| Trace2Env | 93% | 88% | 85% | 0.914 |
| SciWorld | ||||
| Direct Prompting | 85% | 92.5% | 45% | 0.529 |
| Trace2Env | 85% | 87.5% | 60% | 0.706 |
| Variant | Total | Effect |
| Direct Prompting | 55.76 | – |
| Schema only | 62.05 | – |
| + Abstraction (rules, contracts, notes) | 60.69 | |
| + Evidence (evidence turns, demonstrations) | 64.18 | |
| Full worldbook | 64.89 | |
| Abstraction given Evidence |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Env. | Rows/traj. | Action observation | Construction traces and Worldbook |
| Terminal | 354 / 76 | tmux keystrokes terminal screen | 20 own Terminus-2 trajectories on Terminal-Bench 2.0 (334 turns). Worldbook: 15 actions; 29 state fields; 3/12 executable-rule candidates; 18 notes; 334 evidence turns. |
| SWE | 472 / 99 | Coding-agent tools tool result | Benchmark’s own sessions, 5-fold cross-fit per scaffold (2,585 visible transitions). Worldbook: 7–15 actions; 51–137 state fields; 7–17 executable rules; 17–30 notes; 935–1,147 evidence turns. |
| Android | 200 / 92 | UI actions screen-element list | Benchmark’s own trajectories, 5-fold cross-fit per sub-source (789 visible transitions). Worldbook: 6–10 actions; 34–319 state fields; 7–23 executable rules; 10–23 notes; 36–334 evidence turns. |
| Web | 200 / 118 | Playwright MCP calls page snapshot | 50 own WebArena trajectories on the benchmark’s tasks. Worldbook: 9 actions; 227 state fields; 18/28 executable-rule candidates; 23 notes; 501 evidence turns. |
| Food | 150 / 15 | Dietary-profile tools profile or status response | 16 reserved successful EnvScaler trajectories (335 turns). Worldbook: 17 actions; 16 state fields; 6/21 executable-rule candidates; 17 notes; 335 evidence turns. |
| Shopping | 84 / 15 | Cart and inventory tools item list or status response | 3 reserved successful EnvScaler trajectories (55 turns). Worldbook: 10 actions; 4 state fields; 8/10 executable-rule candidates; 13 notes; 55 evidence turns. |
| Env. | Method | Format | Factuality | Consistency | Realism | Quality | Total |
| Terminal | Direct Prompting | 66.2 | 41.7 | 65.2 | 52.8 | 52.9 | 55.76 |
| Trace RAG Prompting | 69.3 | 44.4 | 66.7 | 54.9 | 56.7 | 58.42 | |
| Worldbook Prompting | 69.9 | 45.3 | 67.8 | 56.1 | 56.3 | 59.07 | |
| Harness only | 76.0 | 44.8 | 65.2 | 58.8 | 55.4 | 60.06 | |
| Trace2Env | 80.7 | 50.5 | 71.0 | 64.3 | 60.6 | 65.42 | |
| SWE | Direct Prompting | 70.0 | 57.8 | 71.5 | 67.6 | 62.5 | 65.87 |
| Env. | Method | Format | Factuality | Consistency | Realism | Quality | Total |
| Terminal | Direct Prompting | 77.1 | 44.9 | 66.7 | 58.6 | 56.0 | 60.65 |
| Trace RAG Prompting | 78.8 | 47.1 | 67.7 | 59.5 | 56.3 | 61.90 | |
| Worldbook Prompting | 80.7 | 49.6 | 71.2 | 63.9 | 60.8 | 65.24 | |
| Harness only | 78.9 | 46.2 | 67.1 | 60.5 | 57.6 | 62.06 | |
| Trace2Env | 80.1 | 48.9 | 68.9 | 61.9 | 60.9 | 64.12 | |
| SWE | Direct Prompting | 68.0 | 55.2 | 68.6 | 66.0 | 59.7 | 63.50 |
| Env. | Backbone | T2E DP | T2E H | H DP |
| Terminal | GPT-5.6-Sol | [7.18, 12.15] | [3.17, 7.66] | [1.93, 6.75] |
| DeepSeek-V4.1-Flash | [1.13, 5.74] | [ , 4.53] | [ , 3.80] | |
| SWE | GPT-5.6-Sol | [2.98, 6.54] | [ , 2.17] | [2.57, 5.63] |
| DeepSeek-V4.1-Flash | [5.00, 8.60] | [2.53, 5.75] | [1.24, 4.37] | |
| Android | GPT-5.6-Sol | [0.02, 4.63] | [0.05, 5.00] | [ , 1.18] |
| DeepSeek-V4.1-Flash | [5.42, 11.14] | [2.51, 8.36] | [0.69, 5.10] |
| Env. | Stratum | Rows | GPT-5.6-Sol | DeepSeek-V4.1-Flash |
| Terminal | task not among construction tasks | 312 | [ , ] | [ , ] |
| Terminal | task among construction tasks | 42 | ||
| Web | not linked to a construction task | 151 | [ , ] | [ , ] |
| Web | linked (same task or template) | 49 | ||
| Android † | app covered by another fold | 73 | [ , ] | [ , ] |
| Android † | app in no other fold | 127 |
| System | Worldbook | State | Loop | Knowledge shown | Total | Effect (vs. reference) |
| Direct Prompting | none | no | no | none | 55.76 | – |
| Trace RAG Prompting | none | no | no | top-5 raw turns in the prompt | 58.42 | (vs. Direct Prompting) |
| Worldbook Prompting | full | no | no | worldbook block in the prompt | 59.07 | (vs. Direct Prompting) |
| Harness only | none | no | yes | none | 60.14 | (vs. Direct Prompting) |
| Schema only | schemas | yes | yes | none | 62.05 | (vs. Harness only) |
| + Abstraction | schemas + abstraction | yes | yes | rules, contracts, notes | 60.69 | (vs. Schema only) |
| System | Score | Calls | Cost/row (\downarrow$ | Latency (s) |
| Direct Prompting | 55.76 | 1.0 | 0.04 | 26.0 |
| Trace RAG Prompting | 58.42 | 1.0 | 0.051 | 24.7 |
| Worldbook Prompting | 59.07 | 1.0 | 0.051 | 34.7 |
| Harness only | 60.14 | 1.4 | 0.036 | 17.3 |
| Trace2Env without the adaptive loop | 63.09 | 2.4 | 0.054 | 20.4 |
| Trace2Env without state tracking | 63.91 | 3.3 | 0.144 | 27.9 |
| Env. | tokens: DP / RAG / WB | Harness only | Trace2Env | Tool calls / record |
| Terminal | 12k / 15k / 19k | 0.01, 26 s | 0.03, 48 s | 1.3 / 4.1 |
| SWE | 33k / 36k / 39k | 0.18, 14 s | 0.43, 25 s | 1.2 / 3.2 |
| Android | 20k / 24k / 30k | 0.22, 33 s | 0.59, 49 s | 1.9 / 4.6 |
| Web | 25k / 29k / 35k | 0.26, 27 s | 0.68, 45 s | 2.3 / 5.1 |
| Artifact | Content and use |
| Schemas | |
| Action schema | Action types with their argument names and types; a supplied action is canonicalized against it. |
| State schema | Declared state variables with type and mutability; every state effect is checked against it. |
| Grounded evidence | |
| Evidence records | One per aligned transition: the action, its outcome, extracted facts and candidate effects with uncertainty, and the recorded observation verbatim. |
| Demonstrations | Up to three transitions per action type with their observed before/after facts. |
| Setting | Value |
| Reconstruction | |
| Confidence cap for ambiguous transitions | 0.5 |
| Minimum confidence to compile a rule | 0.55 |
| Demonstrations per action type | 3 |
| Simulation | |
| Trust policy | Reference: model-proposed effects accepted after schema and state-constraint checks |
| Case | Required information | Direct Prompting | Trace2Env |
| 1. Heredoc echo | Terminal continuation-prompt behavior not yet observed in the episode | Begins partway through the script and omits much of the heredoc interaction | Retrieves heredoc turns from other tasks together with the induced continuation-prompt note |
| 2. pip install pandas | Package state accumulated from earlier installation output | Predicts pandas alone and misses its remaining dependencies | Uses tracked world.package_versions together with retrieved installation evidence |
| 3. ALFWorld placement | An unsupported put is a no-op, and the box remains in inventory | Accepts the placement and continues from a false state | Preserves the failed transition and state, then provides inventory and help behavior that enables recovery |
| Field | Content |
| official_input | the benchmark’s turn messages for this record, verbatim and complete: every earlier turn (action and real observation) and the current turn’s action; each earlier turn carries the memory id the agent may cite |
| action | the current action, normalized (type and typed arguments) |
| context | caller-supplied context (the current input as the benchmark presents it) |
| state_summary | a bounded summary of the tracked session state |
| state_fields | the declared state fields and their types |
| retrieved | worldbook items selected for the action, each with its Applicability Gate label: the action’s specification, eligible rules with whether they apply now, the applicable rule if any, invariants, notes, renderer contracts with examples, demonstrations |
| Tool | Description |
| list_actions | List the environment’s known actions with their aliases and argument names. |
| inspect_action | Retrieve an action’s contract, the rules eligible for it now, its observation contracts, notes, and demonstrations. |
| search_knowledge | Full-text search over reconstructed knowledge: rules, notes, demonstrations, evidence turns, and actions. |
| read_evidence | Read one raw turn of the episodes the worldbook was built from: its action and an exact span of its real observation, with the neighbouring turns’ actions. |
| read_state | Read the current session state at a dotted path, or the whole state. |
| state_schema | Describe declared state fields (types, mutability, descriptions), optionally under a path prefix. |
| Field | Description |
| effects | typed state changes on declared paths ( set , delete , increment , decrement , append , merge , remove , schedule ) |
| outcome | success , failure , partial , or unknown , as the environment would classify it |
| observation | the exact observation text the environment returns |
| rule_ids | rules whose effects are applied verbatim |
| citations | identifiers of the retrieved artifacts relied on (rules, notes, evidence turns, memory entries) |
| uncertainty | what could not be established from the workspace |