Long-running autonomous agents must reuse accumulated reasoning experience without allowing explicit historical memory and LLM context to grow indefinitely. However, existing memory mechanisms mainly retrieve, summarize, or compress past content and do not directly learn when particular kinds of thinking should be activated or discover new thinking knowledge from temporally dispersed experiences. This paper proposes a situation-conditioned thinking memory framework that transforms historical reasoning experience into a lightweight policy for predicting what should be thought about in the current situation, while leaving detailed reasoning to a large language model. Situations may represent temporal or spatiotemporal evolution rather than only current states. Temporary experiences are also periodically analyzed across multiple independent episodes to identify repeated long-range regularities, which are consolidated into new thinking knowledge and further internalized by the lightweight policy. Experiments show that the learned policy achieves 1.000 F1 on temporal-rule generalization, improves DeepSeek reasoning F1 from 0.789 to 0.868, reduces online processing time from 0.3636 ms to 0.0382 ms per query at 30,000 historical situations, and reaches 1.000 relation-discovery F1 and future-thinking accuracy after sufficient repeated cross-experience evidence.
Fig. 1: Storage as explicit historical experience grows. Sequence-RAG grows with the number of stored histories, while the consolidated thinking policy remains fixed.
Fig. 2: Online cost under growing history size. Sequence-RAG must search a larger explicit memory, whereas the learned policy performs fixed-size inference.
Experiment
Method
Micro-F1
Exp. 2
Proposed Transformer
1.000±0.000
Summary-RAG
0.566±0.045
RF-Summary
0.383±0.033
Exp. 3
Proposed Sequence
1.000±0.000
Current-State-Only
0.711±0.044
TABLE II: Mechanism verification for temporal-order generalization and situation evolution.
Method
Factor Precision
Factor Recall
Factor F1
Decision Accuracy
Avg. Tokens
DeepSeek-Only
0.704±0.055
1.000±0.000
0.789±0.038
0.484±0.115
719.61±0.98
RAG+DeepSeek
0.835±0.039
1.000±0.000
0.877±0.029
0.559±0.106
727.94±0.82
Proposed+DeepSeek
0.815±0.048
1.000±0.000
0.868±0.035
0.575±0.140
728.13±1.18
Oracle+DeepSeek
0.835±0.039
1.000±0.000
0.877±0.029
0.559±0.106
727.94±0.82
TABLE III: End-to-end reasoning with deepseek-v4-flash . Values are mean ± standard deviation over five seeds.
Fig. 3: Factor F1 in the end-to-end DeepSeek experiment. Learned thinking guidance improves over unguided DeepSeek and approaches the explicit-retrieval and oracle references.
Historical situations
Proposed micro-F1
100
0.777±0.077
250
0.912±0.051
500
0.996±0.005
1,000
0.995±0.006
2,000
0.998±0.003
4,000
0.999±0.001
TABLE V: Rare-pattern learning as historical experience accumulates.
Fig. 4: Learning of a rare thinking pattern as the number of historical situations increases. Error bars show one standard deviation over five seeds.
Repetitions/rule
Candidate Recall
Statistics-Only F1
Proposed+DeepSeek F1
Oracle F1
1
0.000±0.000
0.000±0.000
0.000±0.000
1.000±0.000
2
0.000±0.000
0.000±0.000
0.000±0.000
1.000±0.000
4
0.933±0.149
0.333±0.471
0.333±0.471
1.000±0.000
8
1.000±0.000
1.000±0.000
1.000±0.000
1.000±0.000
At 8 repetitions, No-Consolidation, Unordered, and Local relation F1 are all 0.000±0.000 .
TABLE VI: Cross-experience knowledge discovery. Candidate recall and relation F1 are reported over five seeds. Future-thinking accuracy is evaluated at eight repetitions.
Knowledge source
Future-thinking accuracy
No Consolidation
0.333±0.000
Unordered
0.333±0.000
Local
0.333±0.000
Statistics-Only
1.000±0.000
Proposed+DeepSeek
1.000±0.000
Oracle
1.000±0.000
TABLE VII: Future-thinking activation after cross-experience consolidation at eight repetitions per hidden rule.
Fig. 5: Cross-experience relation discovery as the number of independent supporting episodes increases. Markers with lower opacity show individual random-seed results, while dashed lines show the mean performance. The large variation at four repetitions indicates an intermediate stage in which true relations have begun to emerge but are not yet consolidated consistently across runs.
Hong Su received the MS and PhD degrees, in 2006 and 2022, respectively, from Sichuan University, Chengdu, China. He is currently a researcher of Chengdu University of Information Technology Chengdu, China. His research interests include blockchain, cross-chain and smart contract.
Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents. This gap is reflected in their limited performance on continual learning benchmarks such as ARC-AGI-3, especially when models are evaluated out of the box. Various agent harnesses have been proposed to close this gap, and each commits to a strategy for handling long sequences of observations, i.e., what information to save from the environment and how to load it into model context, a choice we argue is particularly consequential. Existing methods for context management face a significant tradeoff, as preserving more information makes retrieving relevant details less tractable. We propose PRO-LONG, a minimal context management framework built around programmatic memory for LLM agents in long-horizon, exploratory settings. PRO-LONG addresses the tradeoff by keeping a complete, structured interaction log and capitalizing on recent progress in coding agents to search this history efficiently. On the full ARC-AGI-3 public game set, PRO-LONG improves over a base coding agent by an average of 18.0 percentage points across frontier models, and matches or exceeds state-of-the-art specialized harnesses (up to 76.1% pass@1) while using 4.2-5.8x fewer tokens. With Fable 5, PRO-LONG achieves 97.4% best@2 at a total cost of $1,750. Relevant code and logs are available at https://github.com/alexisfox7/PRO-LONG.
Long-context LLMs and Retrieval-Augmented Generation defer state tracking and evidence consolidation to query time, which is brittle when facts evolve and answers depend on latent states. We introduce Unified Memory Agent (UMA) for a one-to-many setting: query-agnostic external memory is constructed once from a stream and reused across multiple future QA sessions. A single policy maintains a structured Memory Bank through CRUD operations and answers using both the Memory Bank and raw context. Task-Stratified GRPO uses the mean reward of QA trajectories branching from each sampled memory state to supervise memory maintenance, while normalizing memory and per-question QA groups separately. We also introduce Ledger-QA, a diagnostic benchmark for long-horizon state tracking over accumulated updates. At the 16k budget, UMA-Generalist achieves the highest average score among compared methods across the test-time-learning and accurate-retrieval benchmarks and transfers to Ledger-QA without task-specific training; UMA-Specialist further improves long-horizon tracking after task adaptation. These results support learned proactive memory management for long-context reasoning.
Kehao Zhang, Shangtong Gui, Sheng Yang +2
1Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS) · University of Chinese Academy of Sciences, Beijing, China · 4Li Auto Inc. +1
Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world. We introduce PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent's decision context. Each belief combines an estimate of the current world state with unresolved task requirements, making explicit what the agent still needs to learn and accomplish. To keep this belief reliable and actionable, PoS validates its consistency and monitors task progress to detect Belief Trapping, where the agent continues to act without making meaningful progress toward the goal. Recovery is then tailored to both the trapping pattern and the type of unresolved task requirement. Experiments on four benchmarks spanning execution and diagnosis show that PoS achieves the highest overall performance on every benchmark with all three LLM backbones. Ablations demonstrate the importance of consistency validation and recovery, while context-scaling experiments show resilience to context growth. Together, these results support belief construction and continual maintenance as a foundation for long-horizon context management beyond history retention and compression.
Yu Luo, Jiamin Jiang, Yimin Zuo +9
Nankai University · Alibaba Group · Tsinghua University