cs.CLOct 5, 2026

Long-Horizon Textual World Modeling through Structured Reasoning

Authors: Fangxin Wang, Xiang Gao, Yuguang Yao, Kaiwen Dong, Nikash Walia, Kamalika Das

Organizations: Intuit AI Research

Abstract

World models must predict how an environment evolves under sequences of actions, enabling agents to compare possible futures and reason about counterfactual actions before acting. Long-horizon prediction is commonly obtained by recursively applying a one-step transition model, but intermediate errors can compound over time. Multi-step dynamics models instead condition on a sequence of future actions and predict their consequences directly, but become harder to learn as horizon grows: the model must track interacting state changes across the trajectory, endpoint supervision provides weak credit assignment, and intermediate predictions can remain plausible while losing information needed for later states. We show that these challenges can be addressed by casting the internal evolution of a multi-step transition as structured reasoning over textual world states: reasoning over sparse state changes reduces the burden of state tracking, a predictive-gain objective rewards the learned state for improving over a matched predictor that conditions on raw history instead, and intermediate predictive rewards supervise each state along the trajectory. Because these intermediate states are explicit textual representations of the world, they provide semantically meaningful targets that can be inspected, scored, and corrected during training. Across ScienceWorld, Jericho, and CEO-Bench, our approach achieves the strongest average long-horizon performance against recursive and non-recursive baselines that condition directly on raw history, with gains increasing at longer horizons. In a controlled counterfactual study, our model is also the only one with statistically significant sensitivity to future actions.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform

    May 13, 2026Feisal Alaswad, Batoul Aljaddouh, Maher Alrahhal +2Latent World ModelsLarge Language Model Agents

  2. Beyond the Next Step: Variable-Length Latent World Models for Long-Horizon Planning

    Jun 19, 2026Tianqi Du, Qi Zhang, Yifei Wang +1Latent World ModelsLong-Horizon Planning