Learning to complete tasks in unfamiliar environments with unknown rules remains a key challenge for LLM agents. Current LLM agents often record their discoveries in prose, which may not provide a compact, explicit account of how the environment works. Inspired by how scientists organize observations into testable, predictive theories, we introduce Schema, an agent harness that organizes learning and action through interactive program induction. The LLM agent decides what to investigate and how to act, expressing its evolving understanding of the environment as executable programs. The harness consists of a persistent program workspace and a small set of interfaces for checking these programs against the interaction history, planning within them, and executing plans under step-by-step verification. Schema raises ARC-AGI-3 RHAE from 58.7% to 99.2% with the same base model, solves 100% of the public DiG-bench games, and reaches the median performance of the top-50 human players on MazeBench. Extensive analysis shows the effectiveness of Schema in unknown mechanism discovery, and ablations confirm the contribution of each component.
Figures & tables
Figure 1: Schema observes the environment, encodes discovered mechanisms as executable programs, and uses them to guide planning and action, refining its understanding through further interaction (top). Schema achieves strong performance across all three benchmarks, outperforming baseline harnesses using the same base models (bottom); dashed lines indicate human reference performance. (a) ARC-AGI-3: RHAE over the 25 public games, Claude Fable 5 in Claude Code and in Schema; the human baseline is 100 by construction. (b) DiG-bench: public games won out of 21, GPT-6 Astra. (c) MazeBench: gems collected, GPT-6 Astra in Codex and in Schema, against the top-50 human median.
Figure 2: Overview of Schema. The LLM agent expresses its understanding as an executable program P in a persistent workspace. The harness checks P against recorded interactions and searches within it using goals and procedures chosen by the agent, without consuming environment actions. Committed plans execute one action at a time under prediction checks. A mismatch interrupts execution and returns new evidence for revising P . The ARC-AGI-3 LS20 level 3 example illustrates an unmodeled color-rotator effect with a simplified prediction: the key is predicted to remain orange but turns blue.
Figure 3: Schema improves completion and action efficiency across model configurations. Results on the 25 public ARC-AGI-3 games: (a) RHAE, (b) percentage of games won, and (c) percentage of games achieving 100% RHAE. Each model configuration is evaluated with the basic harness, coding harness, and Schema.
RHAE (%)
Schema (full)
72.9
prose model
58.8
w/o certification
62.3
w/o planning
59.2
w/o verification
52.4
Table 1: Ablation results.
Figure 4: Trace dynamics. (a) Levels cleared versus cumulative steps on CN04 with program and prose representations, alongside humans. (b) Offline history agreement on AR25 with and without certification, using Schema’s Opus 4.8 main-evaluation run. (c) Levels cleared versus cumulative steps on LS20 with and without planning, alongside humans. (d) Steps taken after a prediction mismatch within the same plan without verification.
Figure 5: Schema solves more games with fewer failed attempts. GPT-6 Astra on the 21 public DiG-bench games: (a) game win rate and (b) total lives lost at each reasoning effort; (c) percentage of levels completed in each game at medium effort, grouped by difficulty tier.
Figure 6: Token cost on DiG-bench. Average cost (a) per cleared level and (b) by level index, including input and output tokens at OpenAI’s official API rates ( OpenAI, 2026e ) .
Figure 7: MazeBench. The full world of 256 rooms, with close-ups of rooms CxG, FxC, and MxJ.
Figure 8: Schema sustains progress over long horizons. (a, b) Gems collected and rooms visited over environment actions; crosses mark the last recorded action of each baseline trajectory, and dotted lines show the top-50 human median. (c) Board-state novelty, with horizontal lines indicating each curve’s mean over the shaded interval after 10,000 actions. (d) Token cost per gem at 23 collected gems.
Figure 9: Knowledge accumulates in a compact, reusable program. (a) Non-comment program lines in Schema and note lines in the prose-model variant. (b) Recorded transitions explained per program line. (c) Call sites per function. (d) Forward prediction accuracy over the trailing 1,000 actions.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: The 25 public ARC-AGI-3 games, each shown at the start of its first level.
RHAE (%)
Model
Cutoff
Release
Public
Semi-private
Gemini 3.1 Pro ( Google DeepMind, 2026 )
Jan 2025
Feb 2026
0.4
GPT-5.4 ( OpenAI, 2026b )
Aug 2025
Mar 2026
0.2
Grok 4.20 ( xAI, 2026 )
Sep 2025
Mar 2026
0.1
Claude Opus 4.7 ( Anthropic, 2026d )
Jan 2026
Apr 2026
0.2
GPT-5.5 ( OpenAI, 2026c )
Dec 2025
Apr 2026
0.4
Appendix
Table 2: Frontier models on ARC-AGI-3, from ARC Prize’s reports and leaderboard as of September 2026 ( ARC Prize Foundation, 2026c ; ARC Prize Foundation, 2026b ; ARC Prize Foundation, 2026a ; ARC Prize Foundation, 2026f ; ARC Prize Foundation, 2026d ) . Model dates follow the cited vendor documentation; empty cells are unreported. † Approximate public score for Fable-class models ( ARC Prize Foundation, 2026e ) . § Parentheses show the provider-adapter harness ( ARC Prize Foundation, 2026g ) .
Win rate (%)
Model
Cutoff
Release
All tiers
Tiers 6–7
Gemini 3.1 Pro ( Google DeepMind, 2026 )
Jan 2025
Feb 2026
16.7
0.0
Qwen 3.6 27B ( Qwen Team, 2026 )
Oct 2025
Apr 2026
1.4
0.0
GPT-5.5 ( OpenAI, 2026c )
Dec 2025
Apr 2026
25.7
10.0
GLM-5.2 ( Z.ai, 2026 )
Mar 2026
Jun 2026
15.3
0.0
Kimi K3 ( Moonshot AI, 2026 )
Jul 2026
22.9
0.0
Appendix
Table 3: Frontier models in the basic harness on all 70 DiG-bench games, as reported by the benchmark authors ( Battleday et al., 2026 ) . Win rates average runs within each game, then games within each set; tiers 6–7 contain 20 games. Each game was solved by at least one first-time human player.
Figure 11: The MazeBench environment. The full world of 256 interconnected rooms, with enlarged views of rooms CxG, FxC, and MxJ. Blue circles and connecting lines locate each room within the world.
With tools
Without tools
Model
Cutoff
Release
Gems (%)
Rooms
Gems (%)
Rooms
Gemini 3.1 Pro ( Google DeepMind, 2026 )
Jan 2025
Feb 2026
1.1
5
0.0
1
GPT-5.4 ( OpenAI, 2026b )
Aug 2025
Mar 2026
0.0
2
GPT-5.5 ( OpenAI, 2026c )
Dec 2025
Apr 2026
0.0
4
Claude Opus 4.8 ( Anthropic, 2026e )
Jan 2026
May 2026
0.0
4
GLM-5.2 ( Z.ai, 2026 )
Mar 2026
Jun 2026
0.0
1
Appendix
Table 4: Frontier models on MazeBench’s ASCII track, from the September 14, 2026 leaderboard ( Pappas and Pappas, 2026b ) . Tools permit Python execution. Gem percentages use each run’s available total (71–100); rooms are out of 256. The human row gives the top-50 median, with gems out of 100; individual entries appear in Table 5 .
Rank
Gems
Rooms
Actions
Hours
Rank
Gems
Rooms
Actions
Hours
1
80
233
115,060
55.99
26
29
104
49,221
95.95
2
80
233
151,663
27.21
27
26
139
40,112
12.31
3
80
233
260,784
102.35
28
26
126
68,310
13.64
4
79
233
63,543
10.82
29
21
174
27,181
22.36
5
79
233
206,116
29.43
30
21
129
26,261
5.04
6
79
233
221,343
72.30
31
21
105
31,332
11.62
Appendix
Table 5: Top-50 human entries used for the MazeBench reference, comprising submissions through September 19, 2026 ( Pappas and Pappas, 2026b ) . Names are omitted; gems and rooms are out of 100 and 256, respectively, and hours denote recorded playtime.
Figure 12: Recorded CN04 frames before growth, after four growth actions, and at level completion. The final arrangement joins the pieces without overlapping their bodies.
Figure 13: The same historical AR25 transition, shown with identical crops. Boxes mark the reflection; dashed lines mark its original position. The turn-15 program predicts the observed one-cell movement exactly, while the turn-39 program predicts two cells.
Figure 14: Recorded LS20 observations before and after the first two actions of the final plan. Enlarged crops show the carried symbol rotating to match the target after the down–up pair. The remaining eleven actions reach the goal.
Figure 15: The first prediction mismatch in the 43-action LS20 batch without verification. Identical crops highlight the station’s location: the prediction retains it, but the observation shows floor. Execution continues through another refill and eventual energy depletion; the rebound after zero is an automatic reset.
Current hypothesis
New practice queries
Consequence
Initial probe
a : yes
Propose a rule involving a .
Ends in a
b , c , d , ab , ba : no
ba rules out the current hypothesis.
Starts and ends in a
aa , aba , ac : yes
ac rules out this positional rule.
More a ’s than b ’s
baa : yes; abba : no; cad : yes
All three agree with the count rule.
Appendix
Table 6: All twelve practice queries in P-3 level 3, grouped chronologically by the program hypothesis in use. Hypothesis descriptions summarize the recorded code edits.
String
dca
a
dbdacb
bdcaaa
dcbdba
acd
bcabca
adbcc
# a − # b
1
1
−1
2
−1
1
0
0
Answer
yes
yes
no
yes
no
yes
no
no
Correct
✓
✓
✓
✓
✓
✓
✓
✓
Appendix
Table 7: The eight scored questions in order. The count difference is shown to make the learned rule easy to check.
Figure 16: Ice-rule discovery and reuse, shown on top-down game-engine views. Arrows trace selected probes in IxG (left), which stop at a wall or upon leaving ice, and the complete 38-action route to the gem in NxE (right).
Opus 4.8
Fable 5
Sol xhigh
Sol max
Game
Levels
Human
RHAE
Actions
RHAE
Actions
RHAE
Actions
RHAE
Actions
AR25 ∗
8
748
100.0
269
100.0
298
100.0
261
100.0
278
BP35
9
651
62.9
1,265
93.5
566
28.8
218 (5/9)
60.9
1,347
CD82
6
171
100.0
121
100.0
144
100.0
116
100.0
86
CN04 ∗
6
789
100.0
479
100.0
324
100.0
318
100.0
241
DC22 ∗
6
1,228
38.9
485 (4/6)
98.7
1,205
100.0
1,018
100.0
814
Appendix
Table 8: Per-game ARC-AGI-3 results of Schema with each base model. Levels is the number of levels in the game and Human the total actions of the human baseline; for each model, RHAE is the game’s score and Actions counts actions through the last cleared level. Parentheses give completion counts for partially cleared games. ∗ The ten games with the most human actions, used for the ablations in Section 4.1 .
Figure 17: Per-game progress on ARC-AGI-3 with Claude Opus 4.8 in Schema. Curves connect cumulative action counts at level completion for Schema and the human baseline. The number in each panel is the game’s RHAE.
Figure 18: Per-game progress on ARC-AGI-3 with Claude Fable 5 in Schema, as in Figure 17 .
Figure 19: Per-game progress on ARC-AGI-3 with GPT-5.6 Sol at extra-high effort in Schema, as in Figure 17 .
Figure 20: Per-game progress on ARC-AGI-3 with GPT-5.6 Sol at maximum effort in Schema, as in Figure 17 .
Medium
High
Max
Games won
Basic harness
8.7 ± 2.5
15.0 ± 1.0
18.0 ± 1.0
Schema
19.3 ± 0.6
19.7 ± 0.6
20.3 ± 0.6
Levels completed
Basic harness
118.0 ± 16.7
156.3 ± 5.5
178.0 ± 5.2
Schema
179.7 ± 4.6
183.7 ± 2.3
187.7 ± 4.6
Appendix
Table 9: DiG-bench with GPT-6 Astra over three runs: mean ± standard deviation of the games won (of 21) and the levels completed (of 193).
Figure 21: DiG-bench with GPT-6 Astra over three runs. (a) Games won and (b) levels completed at each reasoning effort: bars show the mean and error bars one standard deviation. (c) Share of games won in each difficulty tier, over all efforts and runs.
Figure 22: Fraction of levels completed in each DiG-bench game at each reasoning effort, mean and standard deviation over three runs, games ordered by tier.
Figure 23: Room coverage of the four MazeBench runs. Each cell is a room in the same 16×16 grid, with columns and rows A–P (e.g., CxG is column C, row G). Panels show each run’s recorded endpoint. Baselines use the September 14, 2026 ASCII tools leaderboard, as in Figure 8 .
Gem
Room
(x,y)
Step
Δ
Gem
Room
(x,y)
Step
Δ
1
HxH
(1,3)
100
100
18
ExF
(2,6)
11,173
730
2
GxH
(10,1)
303
203
19
ExC
(13,12)
11,478
305
3
HxF
(1,1)
512
209
20
DxG
(14,2)
12,469
991
4
GxF
(2,9)
749
237
21
IxJ
(7,14)
12,643
174
5
FxE
(11,2)
1,001
252
22
JxL
(14,14)
13,759
1,116
6
FxF
(10,2)
1,116
115
23
IxE
(1,10)
15,700
1,941
Appendix
Table 10: All 33 gems collected by Schema. Read down the left block, then the right. Room-local (x,y) coordinates range from 0 to 15 in the unrotated top view, increasing rightward and downward. Step is the cumulative action count; Δ counts actions since the previous gem (from the start for gem 1), including exploration elsewhere.