Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects underspecified. It compiles this rulebook into executable code for prediction and planning. Through a continual loop of observation, reflection, rule revision, compilation, and verification, the agent uses prediction errors to refine both the rulebook and its code. Updated code is accepted only when the LLM judges it faithful to the rulebook and cell-exact replay reproduces the observed transitions. We investigate this process as a model-based route to recursive self-improvement (RSI): the agent autonomously explores the environment, revises its world model, and uses verified updates to guide subsequent interaction and learning, while the underlying LLM remains fixed. A population extension maintains multiple world models in parallel, sharing interaction evidence and using their predictions to guide exploration. On ARC-AGI-3, the single-model agent clears every level of all 25 public games, achieves a mean Relative Human Action Efficiency (RHAE) of 100.0, and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes with different openings, without further LLM calls.
Figures & tables
Figure 1: From a rulebook to an executable world model. Numbered rules map to corresponding branches of a simplified transition function. A planner uses the model to predict a route that changes the badge’s pattern and colour, collects a refill, and reaches the matching door within the move budget.
Figure 2: Continual world-model learning with a frozen LLM. The agent follows a continual loop of observation , reflection , rule revision , compilation , and verification . In this example, an observed badge change contradicts a mirror hypothesis. Reflection yields a clockwise-rotation rule, which is recorded in the rulebook and compiled into code. Verification checks the executable against the interaction history and the rulebook. The accepted model then guides further interaction, producing evidence for the next revision. Here Ht , ψt , and θt denote the interaction record, rulebook, and executable after t real actions; πϕ denotes the LLM’s output distribution with fixed parameters ϕ .
Figure 3: A graphical model of recursive self-improvement in Code as Model. The underlying environment state st produces observation ot , which enters the external learning state Zt=(Ht,ψt,θt) . The agent selects at from Zt , and the environment advances to st+1 . The next observation ot+1 , together with Zt and at , determines the information available for updating Zt+1 . Rectangles contain the history, rulebook, and executable; white circles denote hidden states and grey circles observations and actions. P selects actions, and U updates the rulebook–executable pair. Ellipses indicate continued interaction. Completion feedback also enters the interaction history but is omitted from the diagram for clarity.
Figure 4: An ARC-AGI-3 realisation of a rulebook and its executable. Left: an observed transition from game sc25 . Upper right: the final rulebook in world_model.md , whose names are also used to annotate the frame. Lower right: an excerpt from world_model_engine.py implementing the numbered rules. The executable is checked by cell-exact replay before it is used for planning.
Figure 5: Three of the eleven rulebook versions from one run on game tr87 , each paired with the observation that prompted the revision. (1) Before the first action, the rulebook records two competing hypotheses about glyph orientation. (2) Interaction establishes rotation-invariant letter identity and replaces pixel coordinates with a glyph-box rule. (3) Evidence from later levels generalises and reverses the cursor rule. Red marks a claim retracted by a later version and green marks the statement that resolves it.
Figure 6: The 25 public ARC-AGI-3 games.
Game
L
Continual Harness
Dream- Team
OPINE- World
NOOA
baseline1
Ours
ar25
8
17.4
100.0
100.0
100.0
100.0
100
bp35
9
1.1
0.8
2.5
46.2
100.0
100
cd82
6
0.0
79.0
100.0
100.0
100.0
100
cn04
6
46.4
0.0
100.0
100.0
100.0
100
dc22
6
18.2
0.0
82.4
64.7
100.0
100
ft09
6
70.1
100.0
100.0
100.0
100.0
100
Table 1: Per-game RHAE on the 25 public ARC-AGI-3 games. L is the number of levels in the game. RHAE is the level-weighted human-relative efficiency of Eq. ( 36 ). All baseline numbers come from the official ARC Prize scorecards.
Game
Human
Continual Harness
Dream- Team
OPINE- World
NOOA
baseline1
Ours
Ours / Human
ar25
748
687 †
479
381
274
271
253
0.34
bp35
651
274 †
647 †
512 †
623 †
532
429
0.66
cd82
171
275 †
282
161
115
93
82
0.48
cn04
789
1,792 †
444 †
264
441
201
201
0.25
dc22
1,228
1,097 †
355 †
1,485
841 †
898
643
0.52
ft09
208
270
82
112
84
85
75
0.36
Table 2: Per-game action counts. Results are from the same runs and scorecards as Table 1 . The bar shows our count as a fraction of the human count, so a full track would indicate parity. A dagger marks a system that left at least one level of the game uncleared. Its count is therefore what the run spent rather than what solving the game requires. Totals marked with a dagger are not comparable as solution costs.
Actions
Agent turns
Game
Band
Human actions per level
with rulebook
without rulebook
with rulebook
without rulebook
ft09
Easy
35
75
79
132
174
ka59
Medium
104
341
382
298
310
cn04
Hard
132
201
216
248
346
Table 3: Ablating the rulebook, one game per difficulty band. Difficulty is the human action baseline divided by the level count.
Actions
Level
Human baseline
N=1
N=2 population
N=2 vs N=1
L1
71
42
26
0.62×
L2
119
60
58
0.97×
L3
183
98
76
0.78×
L4
98
84
50
0.60×
L5
368
245
106
0.43×
Table 4: Population extension, level by level on wa30 .
Figure 7: How the Pong game works. The agent controls the right-hand paddle and the opponent controls the left-hand paddle. Each player moves its paddle up or down to intercept the ball and return it towards the other side, scoring a point when the opponent fails to return it. The three frames illustrate one return: the ball approaches the agent’s paddle, makes contact, and bounces back towards the opponent.
Method
Approach
Score
Learning cost
(frames)
GDI [ 10 ]
Model-free RL, data-distribution iteration
21
2.0×108
LASER [ 34 ]
Model-free RL, actor–critic with replay
21
2.0×108
IMPALA [ 9 ]
Model-free RL, distributed actor–critic
20.98
2.0×108
Rainbow [ 15 ]
Model-free RL, distributional DQN with replay
20.9
2.0×108
EfficientZero [ 46 ]
Model-based RL, latent dynamics and MCTS
20.1
4.0×105
Table 5: Pong score. Following the evaluation metric used in reinforcement learning experiments, episode score is defined as the agent’s points minus the opponent’s, ranging from −21 to +21 . Our score is averaged over the three episodes shown in Figure 8 , while the published baselines follow their respective evaluation protocols.
Figure 8: Pong episode scores. The three green curves show the learned controller under three different openings. Each reaches +21 without conceding a point. The same LLM without a world model finishes with a score of −19 , and the random-policy episode shown here finishes at −20 . The horizontal axis counts actions taken within each episode.
Figure 9: Three rulebook versions from one Pong learning session show how observed ball returns drive two revisions of the bounce rule shared by both paddles. (1) The initial four-band hypothesis excludes horizontal returns. (2) A horizontal return from a centre hit refutes this hypothesis and prompts a centred-deflection formula that permits vy=0 . (3) Further contacts reveal that the formula overestimates vertical return speeds, leading to an asymmetric band map with ∣vy∣≤2 . Contact offset is the ball’s top row minus the paddle’s top row; vy denotes vertical velocity in pixels per emulator frame. Red marks claims revised later, and green marks new evidence or revised rules. Dashed white trails show recorded ball positions before contact, while solid white trails show positions after contact; step labels identify the supporting observations.