Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects underspecified. It compiles this rulebook into executable code for prediction and planning. Through a continual loop of observation, reflection, rule revision, compilation, and verification, the agent uses prediction errors to refine both the rulebook and its code. Updated code is accepted only when the LLM judges it faithful to the rulebook and cell-exact replay reproduces the observed transitions. We investigate this process as a model-based route to recursive self-improvement (RSI): the agent autonomously explores the environment, revises its world model, and uses verified updates to guide subsequent interaction and learning, while the underlying LLM remains fixed. A population extension maintains multiple world models in parallel, sharing interaction evidence and using their predictions to guide exploration. On ARC-AGI-3, the single-model agent clears every level of all 25 public games, achieves a mean Relative Human Action Efficiency (RHAE) of 100.0, and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes with different openings, without further LLM calls.
Figures & tables
Figure 1: From a rulebook to an executable world model. Numbered rules map to corresponding branches of a simplified transition function. A planner uses the model to predict a route that changes the badge’s pattern and colour, collects a refill, and reaches the matching door within the move budget.
Figure 2: Continual world-model learning with a frozen LLM. The agent follows a continual loop of observation , reflection , rule revision , compilation , and verification . In this example, an observed badge change contradicts a mirror hypothesis. Reflection yields a clockwise-rotation rule, which is recorded in the rulebook and compiled into code. Verification checks the executable against the interaction history and the rulebook. The accepted model then guides further interaction, producing evidence for the next revision. Here Ht , ψt , and θt denote the interaction record, rulebook, and executable after t real actions; πϕ denotes the LLM’s output distribution with fixed parameters ϕ .
Figure 3: A graphical model of recursive self-improvement in Code as Model. The underlying environment state st produces observation ot , which enters the external learning state Zt=(Ht,ψt,θt) . The agent selects at from Zt , and the environment advances to st+1 . The next observation ot+1 , together with Zt and at , determines the information available for updating Zt+1 . Rectangles contain the history, rulebook, and executable; white circles denote hidden states and grey circles observations and actions. P selects actions, and U updates the rulebook–executable pair. Ellipses indicate continued interaction. Completion feedback also enters the interaction history but is omitted from the diagram for clarity.
Figure 4: An ARC-AGI-3 realisation of a rulebook and its executable. Left: an observed transition from game sc25 . Upper right: the final rulebook in world_model.md , whose names are also used to annotate the frame. Lower right: an excerpt from world_model_engine.py implementing the numbered rules. The executable is checked by cell-exact replay before it is used for planning.
Figure 5: Three of the eleven rulebook versions from one run on game tr87 , each paired with the observation that prompted the revision. (1) Before the first action, the rulebook records two competing hypotheses about glyph orientation. (2) Interaction establishes rotation-invariant letter identity and replaces pixel coordinates with a glyph-box rule. (3) Evidence from later levels generalises and reverses the cursor rule. Red marks a claim retracted by a later version and green marks the statement that resolves it.
Figure 6: The 25 public ARC-AGI-3 games.
Game
L
Continual Harness
Dream- Team
OPINE- World
NOOA
baseline1
Ours
ar25
8
17.4
100.0
100.0
100.0
100.0
100
bp35
9
1.1
0.8
2.5
46.2
100.0
100
cd82
6
0.0
79.0
100.0
100.0
100.0
100
cn04
6
46.4
0.0
100.0
100.0
100.0
100
dc22
6
18.2
0.0
82.4
64.7
100.0
100
ft09
6
70.1
100.0
100.0
100.0
100.0
100
Table 1: Per-game RHAE on the 25 public ARC-AGI-3 games. L is the number of levels in the game. RHAE is the level-weighted human-relative efficiency of Eq. ( 36 ). All baseline numbers come from the official ARC Prize scorecards.
Game
Human
Continual Harness
Dream- Team
OPINE- World
NOOA
baseline1
Ours
Ours / Human
ar25
748
687 †
479
381
274
271
253
0.34
bp35
651
274 †
647 †
512 †
623 †
532
429
0.66
cd82
171
275 †
282
161
115
93
82
0.48
cn04
789
1,792 †
444 †
264
441
201
201
0.25
dc22
1,228
1,097 †
355 †
1,485
841 †
898
643
0.52
ft09
208
270
82
112
84
85
75
0.36
Table 2: Per-game action counts. Results are from the same runs and scorecards as Table 1 . The bar shows our count as a fraction of the human count, so a full track would indicate parity. A dagger marks a system that left at least one level of the game uncleared. Its count is therefore what the run spent rather than what solving the game requires. Totals marked with a dagger are not comparable as solution costs.
Actions
Agent turns
Game
Band
Human actions per level
with rulebook
without rulebook
with rulebook
without rulebook
ft09
Easy
35
75
79
132
174
ka59
Medium
104
341
382
298
310
cn04
Hard
132
201
216
248
346
Table 3: Ablating the rulebook, one game per difficulty band. Difficulty is the human action baseline divided by the level count.
Actions
Level
Human baseline
N=1
N=2 population
N=2 vs N=1
L1
71
42
26
0.62×
L2
119
60
58
0.97×
L3
183
98
76
0.78×
L4
98
84
50
0.60×
L5
368
245
106
0.43×
Table 4: Population extension, level by level on wa30 .
Figure 7: How the Pong game works. The agent controls the right-hand paddle and the opponent controls the left-hand paddle. Each player moves its paddle up or down to intercept the ball and return it towards the other side, scoring a point when the opponent fails to return it. The three frames illustrate one return: the ball approaches the agent’s paddle, makes contact, and bounces back towards the opponent.
Method
Approach
Score
Learning cost
(frames)
GDI [ 10 ]
Model-free RL, data-distribution iteration
21
2.0×108
LASER [ 34 ]
Model-free RL, actor–critic with replay
21
2.0×108
IMPALA [ 9 ]
Model-free RL, distributed actor–critic
20.98
2.0×108
Rainbow [ 15 ]
Model-free RL, distributional DQN with replay
20.9
2.0×108
EfficientZero [ 46 ]
Model-based RL, latent dynamics and MCTS
20.1
4.0×105
Table 5: Pong score. Following the evaluation metric used in reinforcement learning experiments, episode score is defined as the agent’s points minus the opponent’s, ranging from −21 to +21 . Our score is averaged over the three episodes shown in Figure 8 , while the published baselines follow their respective evaluation protocols.
Figure 8: Pong episode scores. The three green curves show the learned controller under three different openings. Each reaches +21 without conceding a point. The same LLM without a world model finishes with a score of −19 , and the random-policy episode shown here finishes at −20 . The horizontal axis counts actions taken within each episode.
Figure 9: Three rulebook versions from one Pong learning session show how observed ball returns drive two revisions of the bounce rule shared by both paddles. (1) The initial four-band hypothesis excludes horizontal returns. (2) A horizontal return from a centre hit refutes this hypothesis and prompts a centred-deflection formula that permits vy=0 . (3) Further contacts reveal that the formula overestimates vertical return speeds, leading to an asymmetric band map with ∣vy∣≤2 . Contact offset is the ball’s top row minus the paddle’s top row; vy denotes vertical velocity in pixels per emulator frame. Red marks claims revised later, and green marks new evidence or revised rules. Dashed white trails show recorded ball positions before contact, while solid white trails show positions after contact; step labels identify the supporting observations.
Digital agents must often adapt to new environments whose interfaces, tools, and failure modes are not fully captured by pretrained models. We introduce \textbf{RSIAgent}, a training-free multi-agent framework for recursive self-improvement through autonomous memory construction. RSIAgent coordinates curriculum, actor, and verifier agents to continually explore the environment, validate outcomes, and retain environment-specific knowledge, including reusable causal relationships between actions, conditions, and consequences. It further adopts a \textbf{broad-then-deep} exploration strategy, combining parallel broad recursive self-exploration for discovering diverse environment structures with focused deep self-exploration for uncovering hard cases, hidden constraints, boundary conditions, and previously unknown causal dependencies. The resulting memory is frozen and can be directly reused for downstream tasks without updating model parameters. Experiments on OSWorld-v2 and Agent's Last Exam show that RSIAgent substantially improves strong open-source models, enabling Kimi-K3 and GLM-5.3 to outperform frontier closed-source models including GPT-6.
Sibo Zhu, Shicheng Fan, Xinyue Wang +3
Aether AI · University of California San Diego · ‡Work done during internship in Aether AI +1
World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequences before execution. However, unreliable foresight can be ignored, misused, or even degrade downstream decision-making. In this paper, we introduce WorldEvolver, a self-evolving world model framework that revises its deployment-time context while keeping the downstream agent and all model parameters frozen. WorldEvolver integrates three modules: (i) Episodic Memory, which exploits real action transitions through retrieval-based simulation; (ii) Semantic Memory, which extracts persistent heuristic rules from prediction-observation mismatches; and (iii) Selective Foresight, which filters low-confidence predictions before integrating them into agent reasoning context. We evaluate WorldEvolver on ALFWorld and ScienceWorld, measuring world model prediction accuracy on Word2World and downstream agent success rate on AgentBoard. Extensive experiments show that WorldEvolver achieves the highest prediction accuracy across three backbones and leads other world model baselines on downstream agent success rate, demonstrating that test-time memory revision enhances both predictive fidelity and planning performance.
Xuan Zhang, Wenxuan Zhang, See-Kiong Ng +1
National University of Singapore · Singapore University of Technology and Design · Singapore Management University
Many real-world tasks (e.g., office workflows, scientific experimentation) require LLM agents to interact repeatedly with their environments for context-dependent operations. However, such environments are often not agent-ready. First, information is often scattered and fragmented across the environment. Second, relevant evidence in the environment is often mixed with misleading information and conflicting versions. Third, environments evolve over time, introducing new noise and more challenging tasks. These challenges can substantially degrade performance for state-of-the-art AI agents (e.g., from 83.9% to 57.6%). To address these challenges, we propose Env-Rethink (a system with 27B post-trained model) that supports three main capabilities: (1) It adaptively builds Collection Maps (for organizing related files) and Event Logs (for contextualizing cross-data relationships) to supplement necessary context; (2) It further leverages the post-trained model (through offline trajectory learning) to identify underlying noise issues in the environment; (3) It ultimately evolves environments through virtual event histories that alter environmental states and evidence relationships, producing more tricky ones for further agent improvement. Experiments show that Env-Rethink can effectively improve downstream task performance (with a 15.1 percentage-point increase in mean rubric pass rate across nine models on 30 tasks).
Yukai Wu, Yuanjing Yang, Le Zhou +8
Shanghai Jiao Tong University · Theseus Labs · Tencent Hunyuan