Long-Horizon Textual World Modeling through Structured Reasoning
Organizations: Intuit AI Research
Abstract
World models must predict how an environment evolves under sequences of actions, enabling agents to compare possible futures and reason about counterfactual actions before acting. Long-horizon prediction is commonly obtained by recursively applying a one-step transition model, but intermediate errors can compound over time. Multi-step dynamics models instead condition on a sequence of future actions and predict their consequences directly, but become harder to learn as horizon grows: the model must track interacting state changes across the trajectory, endpoint supervision provides weak credit assignment, and intermediate predictions can remain plausible while losing information needed for later states. We show that these challenges can be addressed by casting the internal evolution of a multi-step transition as structured reasoning over textual world states: reasoning over sparse state changes reduces the burden of state tracking, a predictive-gain objective rewards the learned state for improving over a matched predictor that conditions on raw history instead, and intermediate predictive rewards supervise each state along the trajectory. Because these intermediate states are explicit textual representations of the world, they provide semantically meaningful targets that can be inspected, scored, and corrected during training. Across ScienceWorld, Jericho, and CEO-Bench, our approach achieves the strongest average long-horizon performance against recursive and non-recursive baselines that condition directly on raw history, with gains increasing at longer horizons. In a controlled counterfactual study, our model is also the only one with statistically significant sensitivity to future actions.
Figures & tables
| Method | Avg. | ||||||
| GPT-5.2 † | 92.67 | 86.70 | 81.31 | 78.11 | 72.80 | 69.46 | 80.17 |
| Copy-last | 90.11 | 81.68 | 74.63 | 68.42 | 63.50 | 59.20 | 72.92 |
| TS model † | 83.75 | 78.16 | 72.16 | 70.32 | 68.08 | 66.26 | 73.12 |
| obs-recursive | 92.48 | 87.15 | 84.34 | 78.10 | 74.16 | 70.28 | 81.08 |
| obs-direct | 89.04 | 85.86 | 82.72 | 80.48 | 78.19 | 74.81 | 81.85 |
| latent-recursive | 92.24 | 87.73 | 82.81 | 80.53 | 75.13 | 72.22 | 81.78 |
| Fact-F1 | BLEU | |||||||||||||
| Method | Avg. | Avg. | ||||||||||||
| GPT-5.2 † | 82.13 | 83.55 | 79.22 | 76.34 | 68.67 | 68.21 | 76.35 | 59.67 | 60.27 | 59.29 | 57.26 | 52.71 | 52.67 | 56.98 |
| obs-recursive | 94.45 | 90.53 | 84.70 | 79.42 | 76.76 | 73.66 | 83.25 | 89.16 | 86.01 | 79.91 | 72.44 | 64.98 | 63.25 | 75.96 |
| obs-direct | 97.00 | 95.78 | 91.19 | 87.14 | 82.91 | 79.62 | 88.94 | 93.56 | 93.15 | 89.34 | 83.79 | 79.62 | 76.07 | 85.92 |
| latent-recursive | 93.32 | 90.32 | 88.56 | 86.22 | 86.76 | 85.42 | 88.43 | 90.19 | 90.07 | 87.66 | 86.15 | 85.81 | 85.20 | 87.51 |
| latent-direct | 91.45 | 90.74 | 88.19 | 88.98 | 87.78 | 87.79 | 89.16 | 88.57 | 88.16 | 87.68 | 87.30 | 86.77 | 87.28 | 87.63 |
| BLEU | Score MAE (lower is better) | |||||||||||
| Method | Avg. | Avg. | ||||||||||
| obs-recursive | 38.0 | 25.2 | 20.3 | 18.6 | 15.1 | 23.4 | 0.217 | 1.039 | 1.870 | 3.039 | 4.689 | 2.171 |
| obs-direct | 31.3 | 22.7 | 21.3 | 21.8 | 20.2 | 23.5 | 0.246 | 1.453 | 2.311 | 2.696 | 3.072 | 1.956 |
| latent-recursive | 24.7 | 23.9 | 23.6 | 23.9 | 23.1 | 23.9 | 2.589 | 3.373 | 3.580 | 4.640 | 4.073 | 3.651 |
| latent-direct | 22.4 | 22.6 | 21.7 | 21.8 | 22.0 | 22.1 | 1.927 | 2.357 | 2.567 | 2.459 | 2.737 | 2.409 |
| latent-SR | 25.5 | 25.1 | 24.2 | 23.3 | 24.2 | 24.5 | 1.276 | 1.412 | 1.642 | 2.105 | 2.204 | 1.728 |
| Variant | Avg. | ||||||
| latent-SR (Distill) | 92.59 | 88.18 | 85.21 | 81.19 | 78.58 | 75.21 | 83.49 |
| latent-SR (GRPO, no Distill) | 91.96 | 85.68 | 68.36 | 49.83 | 41.89 | 28.98 | 61.12 |
| latent-SR (Distill+GRPO) | 92.82 | 89.58 | 87.53 | 84.00 | 78.49 | 76.76 | 84.86 |
| predictive gain | 93.01 | 89.37 | 87.96 | 85.05 | 83.14 | 80.98 | 86.58 |
| mid reward | 93.76 | 90.21 | 88.07 | 85.99 | 84.33 | 81.73 | 87.35 |
| Metric | Model | Avg. | ||||||||||
| relevant | latent-recursive | 8.0 | 9.9 | 13.3 | 12.9 | 11.3 | 8.5 | 11.9 | 10.0 | 13.6 | 12.0 | 11.1 |
| latent-SR (Distill) | 6.1 | 5.6 | 5.9 | 6.7 | 5.9 | 5.3 | 5.8 | 4.1 | 5.3 | 3.0 | 5.4 | |
| latent-SR | 3.0 | 2.5 | 1.1 | 2.5 | 2.7 | 1.9 | 0.9 | 2.0 | 0.8 | 0.1 | 1.7 | |
| irrelevant | latent-recursive | 1.2 | 1.1 | 1.4 | 0.2 | 0.4 | 0.7 | 1.0 | 0.2 | 0.3 | 0.4 | 0.1 |
| latent-SR (Distill) | 0.0 | 0.5 | 0.6 | 0.0 | 0.6 | 1.1 | 1.1 | 0.1 | 0.5 | 0.8 | 0.3 | |
| latent-SR | 0.1 | 0.1 | 0.5 | 0.0 | 0.4 | 0.3 | 1.2 | 0.3 | 0.4 | 0.2 | 0.1 |
| Method | Adj.M (cl.) | Sign-test | Perm. |
| obs-recursive | — | ||
| obs-direct | — | ||
| Chronos-2 (FT) | |||
| latent-SR |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Environment | #Train traj. | #Test traj. | #Train ex. | #Test targets | Avg. traj. len. | Avg. obs. tokens |
| ScienceWorld | 540 | 60 | 17,573 | 298 | 32.5 | 40 |
| Jericho | — | 28 a | 4,131 | 415 | 28.6 | 33 |
| CEO-Bench | 165 | 67 b | 4,740 | 128 | 52 | 14,500 c |
| Env. | Stage | LR | Batch | Epochs | GRPO | KL | Max length |
| ScienceWorld | SFT / Distill | 1e-4 | 32–64 | 3 | — | n/a | 8,192–32,768 |
| ScienceWorld | GRPO | 1e-5 | 16–32 | 3 | 4 | 0–0.01 | 2,048 resp.; prompt up to 32,768 |
| Jericho | SFT / Distill | 1e-5 | 32 | 3 | — | n/a | 8,192 |
| Jericho | GRPO | 1e-6 | 16 | 3 | 4 | 0 | 8,192 / 2,048 |
| CEO-Bench | SFT / Distill | 1e-4 | 32 | 3 | — | n/a | 32,768 |
| CEO-Bench | GRPO | 1e-5 | 16 | 3 | 4 | 0 | 32,768 / 2,048 |
| Architecture | Training | h1 | h3 | h5 | h7 | h9 | h10 | avg |
| TS baselines (numeric only) | ||||||||
| persistence | Copy-last-value | 0.9011 | 0.8168 | 0.7463 | 0.6842 | 0.6350 | 0.5920 | 0.7292 |
| Chronos-2 | 4v + 27 action cov | 0.8375 | 0.7816 | 0.7216 | 0.7032 | 0.6808 | 0.6626 | 0.7312 |
| TimesFM 3.0 | 20v multivariate | 0.8869 | 0.8160 | 0.7319 | 0.6564 | 0.6170 | 0.5756 | 0.7140 |
| Chronos-2 | 4v, no covariates | 0.8302 | 0.7796 | 0.7135 | 0.6865 | 0.6536 | 0.6363 | 0.7166 |
| Chronos-2 | 20v, no covariates | 0.8184 | 0.7748 | 0.7118 | 0.6929 | 0.6568 | 0.6371 | 0.7153 |
| Fact-F1 | BLEU | |||||||||||||
| Variant | Avg. | Avg. | ||||||||||||
| latent-SR (Distill) | 91.39 | 90.99 | 88.82 | 90.04 | 87.23 | 85.99 | 89.08 | 84.22 | 86.15 | 81.51 | 82.13 | 79.37 | 77.23 | 81.77 |
| latent-SR (Distill+GRPO) | 91.87 | 93.31 | 91.30 | 89.39 | 90.08 | 87.23 | 90.53 | 91.17 | 91.46 | 91.18 | 89.44 | 88.53 | 86.98 | 89.79 |
| predictive gain | 95.79 | 93.60 | 92.84 | 91.36 | 89.98 | 87.91 | 91.91 | 93.46 | 92.29 | 93.00 | 91.49 | 89.21 | 87.16 | 91.10 |
| latent-SR (Ours) | 96.36 | 93.01 | 92.53 | 93.09 | 90.27 | 89.38 | 92.44 | 94.09 | 93.49 | 92.85 | 93.21 | 91.13 | 89.48 | 92.37 |
| Training stage | Transition | Avg. | Gap | ||||||
| Raw history (obs) | non-recursive (direct) | 89.04 | 85.86 | 82.72 | 80.48 | 78.19 | 74.81 | 81.85 | |
| recursive | 92.48 | 87.15 | 84.34 | 78.10 | 74.16 | 70.28 | 81.08 | ||
| Distill (SFT) | non-recursive (SR) | 92.59 | 88.18 | 85.21 | 81.19 | 78.58 | 75.21 | 83.49 | |
| recursive | 92.33 | 87.10 | 81.74 | 78.19 | 74.50 | 73.26 | 81.19 | ||
| Distill + trajectory RL | non-recursive (SR) | 92.82 | 89.58 | 87.53 | 84.00 | 78.49 | 76.76 | 84.86 | |
| recursive | 92.24 | 87.73 | 82.81 | 80.53 | 75.13 | 72.22 | 81.78 |
| Method | Adj.M (cl.) | Sign-test (cl.) | Perm. (cl.) | Median Adj.M | Sign-test |
| obs-recursive | |||||
| obs-direct | |||||
| Chronos-2 (27act, FT) | |||||
| latent-SR (Distill) | |||||
| latent-SR (Distill+GRPO) | |||||
| latent-SR (+ predictive gain) |
| Candidate | Avg. | ||||||
| Numeric forecasting baselines | |||||||
| Copy-last ⋆ | 90.11 | 81.68 | 74.63 | 68.42 | 63.50 | 59.20 | 72.92 |
| Chronos-2 (4 targets + 27 action covariates) ⋆ | 83.75 | 78.16 | 72.16 | 70.32 | 68.08 | 66.26 | 73.12 |
| TimesFM 3.0 (20 targets) | 88.69 | 81.60 | 73.19 | 65.64 | 61.70 | 57.56 | 71.40 |
| Chronos-2 (4 targets, no covariates) | 83.02 | 77.96 | 71.35 | 68.65 | 65.36 | 63.63 | 71.66 |
| Chronos-2 (20 targets, no covariates) | 81.84 | 77.48 | 71.18 | 69.29 | 65.68 | 63.71 | 71.53 |
| Branch | |||
| Continue as before | |||
| Ground truth | $809,362 | $808,968 | $806,608 |
| latent-SR | $810,839 | $807,989 | $805,080 |
| Chronos-2 | $815,059 | $818,769 | $849,645 |
| Aggressive growth | |||
| Ground truth | $761,984 | $626,447 | $250,289 |
| Raw-history | Latent-based | |
| J1: Retaining an earlier room state. enchanter_step33_k10_s201 ; , . | ||
| Context. Observation describes an oven in the shack; again places the agent inside it. Later observations include failed movements. The forecast window includes take all , D , E , E , put down book , and open oven with lantern ; the recorded directional attempts fail. | ||
| GT. “Opening the oven reveals a loaf of bread.” | ||
| Prediction | “You can’t see any oven here!” | “Opening the oven reveals a loaf of bread.” |
| BLEU | ||
| Interpretation. obs-direct denies the oven’s availability; latent correctly predicts the interaction. Both output the correct score tag, . The outcome is consistent with state retention, but does not establish which history entries the model used. | ||
| Raw-history | Latent-based | |
| S1: Tracking a multi-room route. d10a1a445c15_2_h8 ; , . | ||
| Context. The observed prefix ends in the kitchen. Future actions move to the hallway and greenhouse, pick up a shovel and soil, teleport to the hallway, and traverse doors to the greenhouse and then the hallway. | ||
| GT. “You move through the door to the hallway.” | ||
| Prediction | “You move to the door between kitchen and hallway.” | “You move through the door to the hallway.” |
| Fact-F1 | ||
| Interpretation. obs-direct still references the kitchen–hallway door and does not correctly describe the final traversal; latent predicts the correct destination. | ||
| Field | Ground truth | Raw-history | Latent-based |
| C1: Subscriber depletion. Session 33f8310faa3b ; , . Subscriber stock evolves from to : . | |||
| Cash | |||
| Individual subscribers | |||
| Enterprise seats | |||
| Open issues | |||
| Four-field score | — | ||