Kepler: Auditable World Models for ARC-AGI-3
Organizations: Independent Researcher
Abstract
ARC-AGI-3 evaluates agents in interactive environments whose rules and objectives must be inferred from observation. We present Kepler, an open-source harness that represents hypotheses as executable world models and validates them through retrospective transition checks and conditional prediction checks. Under one frozen Claude Opus 5 configuration, Kepler obtained a server-verified 100.00 RHAE on all 25 public games, with no per-game model selection or score-conditioned reruns. On 181 of 183 completed levels, the final Opus attempt used no more actions than the corresponding median-human baseline. The retained board runs used 8,256 environment actions, of which 7,292 occurred in scored levels. Retained local provider-session records yield 858.0 million tokens, 97.37% cache reads, and a $777.72 cost at September 1, 2026 API list-equivalent rates. We also report three evaluation failures: source-code leakage that produced an invalid perfect run, agents reconstructing a removed harness in a control condition, and autonomous repair masking a broken planner. A single-game observation case study showed that animation frames contained task-relevant information absent from settled text grids. Across the final Claude Opus 5 and GPT-5.6 Sol boards, 48 of 50 game-model cells reached 100. These results indicate that public-set score alone has limited discriminative value and motivate first-attempt, cost-conditioned, and verification-aware reporting.
Figures & tables
| tool | role | game actions |
|---|---|---|
| observe.py | read-only view of the current frame, diffs, per-level baselines | free |
| query.py | compute over state instead of reading it (context thrift) | free |
| backtest.py | replay simulate over the recorded history; prescribed before planning, not mechanically rerun before each commit | none |
| bfs.py | plan inside the certified model (agents may also write their own searches) | free |
| commit.py | the authorized channel to the real game | charged |
| stage | frozen at | mechanism added | composite | |
|---|---|---|---|---|
| Initial loop | not reported | base loop + guarded channel | 97.78 | n/a |
| Added discipline | d37660e + 6c3a38d | escalation ladder, partial predictions, automatic one-shot clean run | 94.01 | |
| Guard corrections | e61f9ad + d422ab3 | score-floor notices replace hard stops; certified post-WIN restart gate | 96.51 | |
| Certify/replay | eac9af2 | mechanical scored attempts ( cleanrun.py ), forensic fixes | 98.86 |
| board | model | composite | games at 100 | notes |
|---|---|---|---|---|
| Initial loop | Claude Opus 5 | 97.78 | 21 | verified card 00c90840 (97.77 server-side; rounding difference) |
| Added discipline | Claude Opus 5 | 94.01 | 18 | published regression (§ 3.1 ) |
| Guard corrections | Claude Opus 5 | 96.51 | 21 | clean-run gate later shown dead (§ 3.2 ) |
| Certify/replay | Claude Opus 5 | 98.86 | 24 | sp80 71.43 (5/6 levels, operator deadline, § 6.3 ) |
| Initial loop | GPT-5.6 Sol (max) | 93.95 | 18 | verified card fd4733fb |
| Guard corrections | GPT-5.6 Sol (max) | 91.46 | not reported |
| grade | score | evidence |
|---|---|---|
| Initial-loop Opus | 97.77 | official card 00c90840; every recorded action re-executed exactly server-side |
| Certify/replay Opus | 98.86 | server replay 96.38 (cards 455e4374 and 1045a78c) after one lf52 trace hit a replay-invalid off-grid click; retained as development evidence, not a release result |
| Certify/replay GPT | 93.99 / 93.97 | card f5f64ae7; all 25 games re-executed exactly; local versus card differs only in rounding |
| Kepler + Opus | 100.00 | server-verified exact : card 91aa2f10; all 25 games re-executed to 100; server and local ledger agree |
| Visual-replay Opus (superseded) | 100.00 | server replay 94.42 (cards 0c6c0990, 2779000b): two traces each contained one off-grid ACTION6 click; transport validation fixed before release |
| Kepler + GPT | 95.97 | card c9f087f3; server 95.9672, matching local to the fourth decimal; all 25 games re-executed |
| system | score | tokens | cost | trace disclosure |
|---|---|---|---|---|
| Tycho + Opus 5 | 100.00 | 1,343M | $2.99k (authors’ API-equiv.) | scorecards and aggregate metrics |
| Retrodict | 99.86 | 659.9M | $654 | full traces |
| baseline1 (ewma_sv 1.6) | 99.0 | not reported | $400 | scorecards |
| VISTA | 100.00 | not reported | not disclosed | replays on arcprize.org |
| AVO (NVIDIA) | 100.00 | not reported | not disclosed | none published |
| Kepler + GPT-5.6 | 95.97 | 2,429.1M (98.0% cached input) | $1,312.14 (list-equiv.) | public final-board ledgers |
| Failure | Relevant check | Remaining blind spot |
|---|---|---|
| Source-reading win | Retained-log access scan | Unlogged or unrecognized access; no filesystem isolation |
| Contaminated control | Tool-call and filesystem evidence | No causal effect estimate from this comparison |
| Broken planner | Direct tool smoke tests | Outcome scores alone do not reveal substitution |
| Prediction mismatch | Conditional plan interruption | Missing predictions can permit execution |
| Invalid replay input | Coordinate validation and server replay | Certification checks are branch-dependent |
| Incomplete cost record | Separate action and usage ledgers | Final replay excludes prior discovery; private provider records required |