Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?
Organizations: National University of Singapore
Abstract
Learning from experience is essential for LLM agents to adapt to unfamiliar and dynmaic environments. Evaluating this ability is therefore important for understanding how effectively agents acquire and use new knowledge. Existing benchmarks have sought to evaluate this ability, but they primarily evaluate tasks whose rules are provided in the instructions or already familiar to pretrained models, making it difficult to distinguish learning from interactions from reasoning with existing knowledge. To address this, we introduce Learn2Play Bench, a benchmark of newly designed text-based games, whose rules are novel or counterintuitive, requiring agents to acquire knowledge through interaction rather than rely solely on pretrained knowledge. These games provide reproducible feedback and automatic scoring, enabling controlled evaluation of learning across repeated attempts. We also vary game instances to test whether agents can apply what they have learned to new situations. Therefore, we evaluate how backbone models, self-evolving methods, and agent harnesses affect agents' learning ability, revealing three findings: (1) Experience retention: Retaining complete records of actions and feedback can support more effective learning than summarizing these experiences into rules or strategies. (2) Human agent gap: Top-performing human players achieve higher peak scores than the evaluated agents. Human explore more varied strategies, and repeat actions less. (3) Harness matters: With the backbone fixed, changing the harness can improve performance while reducing estimated inference cost. Together, these findings provide insights into how LLM agents learn from experience and suggest directions for future work to improve their learning ability. Project website: https://liushiliushi.github.io/learn2play-bench-website/
Figures & tables
| Harness | Max | Mean | LG | LS |
|---|---|---|---|---|
| Claude Opus 5 | ||||
| OpenCode | 74.5 | 53.8 | +21.3 | +5.13 |
| Claude Code | 81.8 | 62.3 | +27.1 | +6.59 |
| GPT-5.6-SOL | ||||
| OpenCode | 57.0 | 41.1 | +14.5 | +3.71 |
| Codex | 74.3 | 54.8 | +22.2 | +5.66 |
| Agent | Control | Re-shuf. | LS drop |
|---|---|---|---|
| OpenCode backbones | |||
| Claude Opus 5 | 4.10 | 3.51 | 0.59 |
| Qwen 3.8 Max | 3.16 | 3.57 | -0.41 |
| Gemini 3.5 Flash | 4.50 | 2.76 | 1.73 |
| Kimi K3 | 4.19 | 3.31 | 0.88 |
| Gemini 3.7 Flash | 4.49 | 2.34 | 2.15 |
Appendix figures & tables67 assets
Supplementary material from the paper’s appendix.
Appendix
| Game | What the player needs to learn |
|---|---|
| Poisoner | Infer taste constraints from accepted and rejected dishes and apply them to new menus. |
| SynergyCorp | Learn which character traits predict preferred, weak, and taboo actions. |
| Haunted Inn | Discover guest–room, adjacency, and count taboos through controlled placements. |
| Roommates | Infer pairwise and nonlinear grouping rules from room-level feedback. |
| Roadside Observatory | Learn social protocols and resource requirements for new encounters. |
| Lost and Found | Learn how to verify ownership without revealing private information. |
| Game | What the player needs to learn |
|---|---|
| PrimordialSoup | Discover reaction recipes and plan synthesis within a resource budget. |
| GemForge | Discover multi-tier recipes and manage limiting ingredients. |
| Catnip | Learn action–target matches and when to exit for the best reward. |
| Dungeon | Remember rooms, keys, hazards, and treasures to refine an escape route. |
| DreamGarden | Infer delayed ecological effects and plan actions across turns. |
| Ecosphere | Infer food-web, competition, and waste dynamics to plan stocking. |
| Game | Seed | Source | Normalized score | |
|---|---|---|---|---|
| PrimordialSoup | 1 | 110 | exact | |
| GemForge | 1 | 791 | exact | |
| Catnip | 1 | 320 | estimated | |
| Dungeon | 2 | 54 | exact | |
| SynergyCorp | 2 | 137 | exact | |
| DreamGarden | 1 | 6300 | estimated |
| System | Model | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Human Top-1 | Human | 84.3 | 53.5 | +21.6 | +7.33 |
| Claude Code | Claude Opus 5 | 81.8 | 62.3 | +27.1 | +6.59 |
| Human Top-3 | Human | 79.1 | 50.1 | +18.9 | +5.54 |
| Human Top-5 | Human | 75.1 | 47.1 | +17.1 | +4.62 |
| OpenCode | Claude Opus 5 | 74.5 | 53.8 | +21.3 | +5.13 |
| Codex | GPT-5.6-SOL | 74.3 | 54.8 | +22.2 | +5.66 |
| Method | Control | Re-shuf. | LS drop |
|---|---|---|---|
| Naive | -0.49 | -0.41 | -0.08 |
| Memory | 3.76 | 2.88 | 0.88 |
| Reflexion | 3.33 | 2.70 | 0.63 |
| ReasoningBank | 2.48 | 2.09 | 0.40 |
| AWM | -0.11 | 0.26 | -0.37 |
| EvoTest | 1.92 | 1.07 | 0.85 |
| System / model | Control LS | Control LS incl. extra |
|---|---|---|
| OpenCode Claude Opus 5 | 4.10 | 4.07 |
| OpenCode Qwen 3.8 Max | 3.16 | 3.34 |
| OpenCode Kimi K3 | 4.19 | 4.33 |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Human Top-1 | 1 | 100.0 | 100.0 | – | – |
| OpenCode + GPT-6 Astra | 3 | ||||
| Codex + GPT-5.6-SOL | 4 | ||||
| Claude Code + Opus 5 | 3.7 | ||||
| OpenCode + Kimi K3 | 4.7 | ||||
| OpenCode + Qwen 3.8 Max | 4 |
| Target | Criterion |
|---|---|
| T1 | Both base reactions: and . |
| T2 | A and F are inert, and one react action fires each feasible base recipe at most once. |
| T3 | React consumes B and sets the catalyst level to . |
| T4 | Catalyst persists through synthesis but is overwritten by the next react action. |
| T5 | Catalyst level at least one enables . |
| T6 | Free does not consume catalyst and takes priority over L2b. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Human Top-1 | 4 | 99.9 | 59.4 | – | – |
| Claude Code + Opus 5 | 3.7 | ||||
| EvoTest + Kimi K3 | 3 | ||||
| OpenCode + Gemini 3.5 Flash | 3 | ||||
| OpenCode + Qwen 3.8 Max | 4 | ||||
| Codex + GPT-5.6-SOL | 4.3 |
| Target | Criterion |
|---|---|
| T1 | Complete five-recipe map from raw materials to tier-1 gems. |
| T2 | Complete four-recipe map from tier-1 to tier-2 gems. |
| T3 | Complete two-recipe map from tier-2 to tier-3 gems. |
| T4 | Every nonrecipe pair consumes its inputs and yields slag. No three-input recipe exists. |
| T5 | Final inventory determines score. Solo firing yields one-point slag. Ash and crystal do not advance to higher tiers. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Claude Code + Opus 5 | 9 | ||||
| Human Top-1 | 8 | 70.3 | 53.5 | – | – |
| OpenCode + GPT-6 Astra | 6.3 | ||||
| EvoTest + Kimi K3 | 3 | ||||
| Codex + GPT-5.6-SOL | 4.3 | ||||
| Reflexion + Kimi K3 | 7.3 |
| Target | Criterion |
|---|---|
| T1 | The six targets occur in a fixed order. |
| T2 | Complete mapping from all six targets to their sensory preferences. |
| T3 | Complete mapping of soothing and chaotic style preferences. |
| T4 | The five influence actions have distinct modalities and styles, and matching them changes influence gain. |
| T5 | Conversion uses target-specific LOW, MED, and HIGH reward thresholds. |
| T6 | Complete breed routing: Ragdoll to cuteness, Bobcat to menace, Black cat to stealth, and Orange to energy. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Claude Code + Opus 5 | 2.3 | ||||
| Human Top-1 | 5 | 100.0 | 67.8 | – | – |
| OpenCode + Opus 5 | 4.7 | ||||
| OpenCode + Opus 5.5 | 3 | ||||
| OpenCode + Gemini 3.5 Flash | 4.7 | ||||
| EvoTest + Kimi K3 | 3 |
| Target | Criterion |
|---|---|
| T1 | Complete 20-room graph, including every visible edge and every directed hidden passage. |
| T2 | Room 14 is the exit and grants a ten-point escape bonus. |
| T3 | Room 16 contains the 25-gold treasure. |
| T4 | The key is in room 17 and is consumed when room 16 is opened. |
| T5 | Complete location map for the charm, oil, boots, and potion. |
| T6 | Complete effect map for those four items. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| OpenCode + Kimi K3 | 5 | ||||
| Human Top-1 | 5 | 91.2 | 44.5 | – | – |
| Claude Code + Opus 5 | 5 | ||||
| OpenCode + Opus 5.5 | 5 | ||||
| OpenCode + Opus 5 | 5 | ||||
| Codex + GPT-5.6-SOL | 5 |
| Target | Criterion |
|---|---|
| T1 | Complete partition of the eight actions into four behavior pairs. Designer-assigned modality names are not required. |
| T2 | Complete mapping of eight informative desk traits and identification of four decoys. |
| T3 | Complete preference ordering over the four behavior groups for each boss type. |
| T4 | The overtime-selfie and buzzwords taboo clears approval and adds eight demerits. |
| T5 | Repeating one action produces two-stage decay, motivating alternation within a pair. |
| T6 | Steal-lunch followed by buzzwords applies the ordered combo bonus. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Claude Code + Opus 5 | 4.7 | ||||
| OpenCode + Opus 5 | 7.7 | ||||
| Human Top-1 | 8 | 84.2 | 78.7 | – | – |
| Reflexion + Opus 5 | 7.3 | ||||
| OpenCode + Kimi K3 | 8.7 | ||||
| OpenCode + Qwen 3.8 Max | 5.7 |
| Target | Criterion |
|---|---|
| T1 | Resonance per turn is . Population is capped at 100 and decays by 20%. For seed 1, Spores and Moths are worth , Nightmares , and Jelly and Weavers 0. Nightmares suppress the Weaver chain. |
| T2 | Eclipse increases Moths but also introduces Nightmares. Throb increases Weavers, Whisper suppresses Nightmares, and Tide feeds only Nightmares. |
| T3 | Throb starts the chain Weavers Jelly Spores. Each ecological link takes one turn, so the final benefit is delayed. |
| T4 | Over 18 turns, combine direct Moth growth, delayed Spore growth, and Nightmare removal. Throb has little immediate value but a large payoff two turns later. |
| T5 | The archived 5,367-point sequence starts with Throb, then interleaves Eclipse and Whisper to keep Moths and Spores near saturation and Nightmares near zero. It never uses Tide. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Human Top-1 | 5 | 85.0 | 59.7 | – | – |
| Claude Code + Opus 5 | 2.3 | ||||
| OpenCode + Opus 5.5 | 2.7 | ||||
| OpenCode + Gemini 3.5 Flash | 4 | ||||
| Reflexion + Opus 5 | 4 | ||||
| Naive + Opus 5 | 3 |
| Target | Criterion |
|---|---|
| T1 | Spend a one-time budget of 150 on stocking, using at most 40 setup actions. Sealing the habitat starts a deterministic 30-day simulation. |
| T2 | Moss, algae, and microbes support the ecosystem but score zero. Herbivores score 4 per animal, subject to saturation. Minnows and newts score 20. |
| T3 | Seed-1 producer layer: moss and algae biomass per unit, each capped at 10 stocked units. |
| T4 | Seed-1 food web and competition: snail and shrimp eat moss, daphnia eats algae, rank order gives shrimp priority over snail, and overgrazed producers collapse. |
| T5 | Predator and keystone dynamics: minnow hunts shrimp only, newt hunts shrimp/snail, prey below refuge 5 cannot be caught, and a single newt culls shrimp enough for snails to survive. |
| T6 | Living animals and corpses create waste that produces toxic gas. Enough decomposers must be stocked to clear it. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Human Top-1 | 5 | 93.8 | 52.5 | – | – |
| OpenCode + Opus 5.5 | 4.3 | ||||
| Claude Code + Opus 5 | 4.3 | ||||
| OpenCode + GPT-6 Astra | 2 | ||||
| Memory + Opus 5 | 3.3 | ||||
| OpenCode + Gemini 3.7 Flash | 4 |
| Target | Criterion |
|---|---|
| T1 | Choose at most three seeds in a deterministic 20-node linear-threshold cascade. Each episode allows one release. Feedback after each wave provides evidence about activation thresholds and incoming connections. |
| T2 | Seed-1 synergy trap: best singles are T/D at 5, the greedy high-single trio reaches only about 8, while the exact optimum reaches 16. |
| T3 | Before release, add or remove seeds to compare candidate seed sets across episodes and identify which nodes contribute additional activation. |
| T4 | Deterministic exact-trio search over the discovered activation frontier, not choosing T/D/Q/L/M/B by single-node strength. |
| T5 | 16-reach triples: EFT activates KQS, then CMR, GH, JO, I, D, L and leaves A/B/N/P. EPT activates KQS, then CM, GHR, JO, I, D, L and leaves A/B/F/N. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Human Top-1 | 1 | 100.0 | 50.0 | – | – |
| EvoTest + Kimi K3 | 1.3 | ||||
| Naive + Kimi K3 | 3 | ||||
| OpenCode + GPT-6 Astra | 3 | ||||
| Naive + GPT-5 Mini | 3.7 | ||||
| Codex + GPT-5.6-SOL | 4.7 |
| Target | Criterion |
|---|---|
| T1 | Eight rooms with visible fixed features. For the evaluated seed 2, rooms 1, 3, 6, and 7 are plain. Room 2 has a shrine, room 4 a window, room 5 is dark, and room 8 has a mirror. |
| T2 | All three hidden taboos for the evaluated seed 2: housing at least three actors haunts every actor room. Adjacent actors and scholars haunt both rooms. Any lantern-bearing guest in the mirror room is haunted. |
| T3 | Clues identify rooms that trigger taboos. Panic then empties the entire connected block of occupied rooms. An empty room stops that spread. |
| T4 | Place a few guests at a time to distinguish the room that triggered a taboo from rooms affected by spreading panic. Use these comparisons to separate adjacency rules from guest-count rules. |
| T5 | Fee-preserving packing under constraints: keep high-fee merchants/actors when legal, turn away or separate conflicts, and use empty gaps only when they save more fee than they cost. |
| T6 | 200-point seating policy on favorable queues: avoid the three taboo triggers, insert firebreaks where needed, and do not sacrifice high-fee guests for untested superstition. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Human Top-1 | 4 | 100.0 | 54.8 | – | – |
| OpenCode + GPT-6 Astra | 4.3 | ||||
| Claude Code + Opus 5 | 2.7 | ||||
| OpenCode + Kimi K3 | 4 | ||||
| Codex + GPT-5.6-SOL | 3.7 | ||||
| OpenCode + Opus 5.5 | 3.7 |
| Target | Criterion |
|---|---|
| T1 | Seven-place seed-1 island map from beach to forest, swamp, cave, ravine, bamboo, and ruins. |
| T2 | Resource/recipe table: wood/flint/vine/water, seed-1 food effects, beach fire, torch, tool, and raft, including mandatory tool and two-vine raft lashing. |
| T3 | Survival and hazard model: 56 segments, five segments per day, hunger drain, day survival points, cave current, swamp mosquitoes/collapse, bamboo falling rocks/night trap, ruins crab, ravine boar, and forest monkey. |
| T4 | True-vs-decoy escape chain: frame at cave, sail at forest, rudder at ruins are real, while compass at bamboo is the decoy that makes raft crafting fail if used as a substitute. |
| T5 | 241-point script: build fire/tool/torch, route around the two lethal traps, neutralize guards, collect all real raft parts while optionally collecting the decoy for information/score, safely stall for day points, and craft the raft at the beach. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| OpenCode + Opus 5.5 | 4.3 | ||||
| Reflexion + Opus 5 | 6 | ||||
| OpenCode + Gemini 3.5 Flash | 5.3 | ||||
| Human Top-1 | 7 | 89.6 | 47.9 | – | – |
| OpenCode + GPT-6 Astra | 2.3 | ||||
| Claude Code + Opus 5 | 8.3 |
| Target | Criterion |
|---|---|
| T1 | Travel-state accounting for distance, time, energy, supplies, cash, rest, bus, safe hitching, and main-road walking. |
| T2 | Offer to help the mechanic before asking for a ride. Asking without first helping wastes the encounter. |
| T3 | Keep energy at least 50. Ask the ranger about weather and the trail before asking for the shortcut. |
| T4 | Hiker protocol: pay supply/cash or forage, then listen for the route clue. |
| T5 | Collector protocol: visit landmark or keep token, then tell story/show token to receive the coded map clue. |
| T6 | Driver protocol and trap avoidance: refuse offer, question detail, then accept only the verified safe ride. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| AWM + Opus 5 | 3.3 | ||||
| OpenCode + GPT-6 Astra | 1.7 | ||||
| Memory + Opus 5 | 4 | ||||
| OpenCode + Opus 5 | 5 | ||||
| Memory + Kimi K3 | 3.7 | ||||
| Naive + Opus 5 | 2 |
| Target | Criterion |
|---|---|
| T1 | Three reshuffled cases per episode with four claimant roles: owner, finder, associate, and mimic. |
| T2 | Inspect the item privately to learn its concealed marker. Opening it publicly or asking questions too early reveals details that later claimants can repeat. |
| T3 | Verify the concealed detail, the claimant’s relationship to the item and its purpose, the loss timeline, and the relevant external record. A matching decoy record is not sufficient. |
| T4 | Resource budgeting: 24 turns, eight privacy tokens, three isolation tokens, and three record checks across all cases. |
| T5 | 1,000-point office protocol: inspect each item privately, verify owner on all independent channels, check the matching record, and return or store without leaking decisive evidence. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Reflexion + Opus 5 | 10 | ||||
| AWM + Kimi K3 | 8.7 | ||||
| Human Top-1 | 7 | 68.4 | 40.7 | – | – |
| ReasoningBank + Opus 5 | 8.7 | ||||
| OpenCode + GPT-6 Astra | 6.3 | ||||
| EvoTest + Opus 5 | 3 |
| Target | Criterion |
|---|---|
| T1 | 21-turn, seven-day shop economy with cash, unit cost, launch stock, stockouts, expired inventory, and regular-customer churn. |
| T2 | Recipe rule: only products at distance 1 from the classic drink are viable, while far variants can sell briefly but fail the adoption chain. |
| T3 | Customer-role inference from clues: Explorer, Echo, Verifier, Broadcaster, and Anchor/regulars have distinct responses to sample, feature, discount, gift, and bundle. |
| T4 | Stage transition chain LAUNCHED SEEDED REPEATED VERIFIED TRENDING, with clean intervals and short verification/diffusion deadlines. |
| T5 | Sampling Explorers starts adoption. Feature promotion works after a natural repeat purchase. A Broadcaster gift works only after verification. Bundles pay off only after the product is trending. Repeated discounts teach customers to expect low prices. |
| T6 | Inventory/novelty control: restock before return and bundle windows, use hold/restore-classic to reduce novelty pressure, and avoid leftover-expiry costs. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Human Top-1 | 2 | 100.0 | 58.0 | – | – |
| Reflexion + GPT-5 Mini | 3 | ||||
| OpenCode + Opus 5.5 | 1 | ||||
| Codex + GPT-5.6-SOL | 4.7 | ||||
| Claude Code + Opus 5 | 4 | ||||
| OpenCode + Gemini 3.5 Flash | 2.3 |
| Target | Criterion |
|---|---|
| T1 | Line and action economy: 20 people, 14 nights, four AP per day, three quarantine cells, severity 1–3, and death after severity . |
| T2 | Hidden-carrier dynamics: two asymptomatic carriers are the only spreaders, reshuffled into left/right halves, with radius , while symptomatic non-carriers do not infect anyone. |
| T3 | Intervention semantics: quarantine blocks both catching and spreading, whereas treatment only reduces current severity and gives one-night immunity after cure. |
| T4 | Carrier diagnosis from spatial evidence, especially separated sick pairs around a healthy center and disappearance of new cases after quarantining that center. |
| T5 | AP triage: quarantine likely carriers early, treat only patients near death, and release/replace false quarantines without exceeding three cells. |
| T6 | For a 100-point containment policy, identify both carriers by day 2–3 and quarantine them before their spread radius grows. Time treatments so that no other villager reaches the death threshold. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| EvoTest + Opus 5 | 3.7 | ||||
| ReasoningBank + Opus 5 | 4.3 | ||||
| Claude Code + Opus 5 | 2.7 | ||||
| OpenCode + Opus 5 | 2.7 | ||||
| Human Top-1 | 5 | 89.6 | 59.2 | – | – |
| OpenCode + Opus 5.5 | 4 |
| Target | Criterion |
|---|---|
| T1 | Per-level action grammar: observe evidence, optionally audit risky records, diagnose source/type/strategy, then patch the root object. |
| T2 | Levels 1–3 mappings: coin/drain uses environment-cover pavement friction, vibrating cup requires auditing residual patch then expiring it, and flickering key requires object physics repair by solidifying the key. |
| T3 | Levels 4–5 mappings: hospital ledger is an identity-record trap solved by binding the wristband, while premature echo is a time-order conflict solved by aligning the acoustic log. |
| T4 | Level 6 ordered record repair: reconcile the ledger before halting archive reindexing. |
| T5 | Level 7 public-record exposure solved by camera desync. |
| T6 | Level 8 identity-anchor pollution solved by reclaiming the name anchor. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Human Top-1 | 3 | 100.0 | 66.8 | – | – |
| OpenCode + Gemini 3.7 Flash | 3.3 | ||||
| OpenCode + DeepSeek V4 Pro | 3.7 | ||||
| Claude Code + Opus 5 | 3.7 | ||||
| OpenCode + GPT-5.6-SOL | 4.3 | ||||
| OpenCode + Opus 5.5 | 2.3 |
| Target | Criterion |
|---|---|
| T1 | Three-banquet inference interface: eight dishes per banquet, one serve action per banquet, score , and only one labeled accept/reject outcome per menu. |
| T2 | The exact five-feature conjunction for the evaluated seed 2: not sweet, spicy, meaty, fragrant, and soft are required. Sourness, saltiness, and temperature are irrelevant. |
| T3 | Wildcard and near-miss handling: each menu contains at least one full match and at least one 4-of-5 bait, so four-feature theories fail. |
| T4 | For each new menu, check the learned combination of required features. Ignore irrelevant traits and distinguish a full match from dishes satisfying only four requirements. |
| T5 | 3/3 regicide: carry the exact five locked features across banquets and serve a full match in every reshuffled menu. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Human Top-1 | 3 | 100.0 | -4.7 | – | – |
| Reflexion + Opus 5 | 3.7 | ||||
| ReasoningBank + Opus 5 | 3.3 | ||||
| Reflexion + GPT-5 Mini | 4.3 | ||||
| Claude Code + Opus 5 | 4 | ||||
| OpenCode + DeepSeek V4 Pro | 4.7 |
| Target | Criterion |
|---|---|
| T1 | Feedback decomposition: one-shot assignment of nine tenants into three triple rooms, then infer pairwise value from room-level harmony sums rather than opaque labels. |
| T2 | Seed-1 hidden five-type ring: Extrovert–Homebody–Foodie–Night Owl–Neat Freak up to rotation/reversal, where adjacent types score , distance-two types score , and same type scores 0. |
| T3 | Nonlinear trio chemistry: for Night-Owl/Extrovert/Homebody and Foodie/Foodie/Neat-Freak, and for Foodie/Neat-Freak/Homebody. |
| T4 | Search the 280 possible assignments for the current tenants. Account for the part of each room’s score that pairwise values alone cannot explain. |
| T5 | 7-point partition: combine the ring pair map with the three trio exceptions and pick all three rooms jointly, not greedily one room at a time. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| OpenCode + Opus 5 | 8.3 | ||||
| OpenCode + GPT-6 Astra | 9 | ||||
| OpenCode + Qwen 3.8 Max | 6.7 | ||||
| OpenCode + Gemini 3.5 Flash | 7.3 | ||||
| OpenCode + Opus 5.5 | 7.7 | ||||
| Claude Code + Opus 5 | 5.3 |
| Target | Criterion |
|---|---|
| T1 | Six-location time-loop map on the Sea-Azure, with the first loop as a forced prologue and later loops retaining player memory. |
| T2 | Culprit evidence channels: oldman/backpacker saw graycoat track the suite owner, caller heard the insurance phone call, dining search finds over-insurance and forged seaworthiness evidence, and graycoat can supply documents. |
| T3 | Role inversion: graycoat is the ex-engineer ally seeking the stern scuttling rig, while suit is the owner staging insurance fraud with an engine-room accomplice. |
| T4 | Gating rules: confronting suit needs hard evidence plus at least two independent culprit clues. Calling coast guard needs evidence and must happen before the radio window closes at segment 24. |
| T5 | Trust and device prerequisites: comfort graycoat only after suit is restrained and after learning his victim story, and disarm only after comfort plus device knowledge. |
| T6 | Failure avoidance: restraining graycoat, accusing without evidence, comforting too early, or solo disarming ends the loop with only milestone credit. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| EvoTest + Opus 5 | 2.7 | ||||
| Claude Code + Opus 5 | 4.3 | ||||
| EvoTest + Kimi K3 | 2.3 | ||||
| Human Top-1 | 4 | 91.6 | 62.0 | – | – |
| Codex + GPT-5.6-SOL | 2 | ||||
| OpenCode + Kimi K3 | 4.7 |
| Target | Criterion |
|---|---|
| T1 | Breeding interface: four visible traits, deterministic offspring, six mate actions, permanent founders, and capacity-limited offspring management. |
| T2 | Antenna dominance: seed-1 strength order has Feathery strongest, then Straight, Hooked, and Curled weakest. |
| T3 | Size law: child size is the half-up average of parent indices, making extreme sizes require matching/extreme donors. |
| T4 | Color law: seed-1 seven-color nontransitive ring Purple–Blue–Green–Yellow–Cyan–Orange–Crimson, with each color beating the next three and no global strongest color. |
| T5 | Pattern law: normally max-complexity wins, but Curled antenna suppresses the child pattern to Plain. |
| T6 | Scoring law: offspring gain by target-match count , novelty bonus, duplicate decay, and first-exact unused-cross bonus. |
| Configuration | First | Max | Mean | LG | LS |
|---|---|---|---|---|---|
| Naive + Kimi K3 | 1 | ||||
| OpenCode + Kimi K3 | 1.7 | ||||
| EvoTest + Kimi K3 | 1.7 | ||||
| EvoTest + Opus 5 | 1.7 | ||||
| Claude Code + Opus 5 | 2 | ||||
| OpenCode + Opus 5 | 2 |
| Target | Criterion |
|---|---|
| T1 | Each episode has a new machine with ternary inputs. What transfers is the method for testing and modeling the machine, not the previous exam answers. |
| T2 | The machine combines four mechanisms: GATE thresholds signed current inputs. MIX sums modulo 3. HOLD maintains a bounded up/down state. ECHO delays its input by one step. |
| T3 | A two-input GATE/MIX component feeds a shared hidden HOLD/ECHO state. One output is a reversible MIX view of that state. The other uses GATE/MIX and may also depend on an additional state component. |
| T4 | Reset between probe sequences. Single-input ramps or pulses, pairs of active inputs, and graded mixed inputs distinguish GATE from MIX and HOLD from ECHO. |
| T5 | Fit both outputs jointly within the 40-probe budget. Fitting each output separately can miss the state they share. |
| T6 | For full exam accuracy, fit a depth-3 machine consistent with the observed transcript and check that all consistent candidates predict the same outputs for the unseen ten-step exam. Submit every output without relying on feedback before the exam ends. |