Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world. We introduce PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent's decision context. Each belief combines an estimate of the current world state with unresolved task requirements, making explicit what the agent still needs to learn and accomplish. To keep this belief reliable and actionable, PoS validates its consistency and monitors task progress to detect Belief Trapping, where the agent continues to act without making meaningful progress toward the goal. Recovery is then tailored to both the trapping pattern and the type of unresolved task requirement. Experiments on four benchmarks spanning execution and diagnosis show that PoS achieves the highest overall performance on every benchmark with all three LLM backbones. Ablations demonstrate the importance of consistency validation and recovery, while context-scaling experiments show resilience to context growth. Together, these results support belief construction and continual maintenance as a foundation for long-horizon context management beyond history retention and compression.
Figures & tables
Figure 1: Performance and motivation of PoS . Left: PoS outperforms the strongest baseline for each backbone across four benchmarks. Right: a real ALFWorld case contrasts implicit belief reconstruction that leads to Belief Trapping with explicit belief maintenance in PoS , which enables consistency validation and recovery toward task completion.
Figure 2: Overview of PoS . Belief Modeling constructs and incrementally updates a task-conditioned Belief from interaction evidence (Section 3.1 ). Belief-Guided Interaction conditions each action on the current Belief and an Active Gap, while the Belief Sentinel validates consistency and records task-relevant progress (Section 3.2 ). Trapping-Aware Recovery estimates Belief health, diagnoses heterogeneous trapping conditions, and composes recovery constraints to restore progress (Section 3.3 ).
Table 1: Main results across four benchmarks. We report overall task success for ALFWorld (ALF) and LOCA-Bench (LOCA), joint accuracy for RCA-100 (RCA), and overall diagnosis accuracy for ClinDiag (Clin), all in %. Best performance is bold. Fine-grained results in Table 9
Figure 3: Belief Trapping and recovery across benchmarks. Left: Percentage of episodes with detected trapping under each backbone. Middle: Distribution of trapping patterns by benchmark. Right: Final task performance over all evaluated cases with Qwen3.7-Plus, comparing PoS ’s factorized recovery with a generic recovery prompt.
Backbone
Method
Joint (%)
Tokens per episode (K)
Task Agent
Belief
Sentinel
Total
Qwen3.7-Plus
Raw Trajectory
24.27
355.70
0.00
0.00
355.70
PoS (Ours)
38.83
281.20
696.69
822.24
1800.13
- Consistency Validation
31.07
299.82
733.41
125.30
1158.53
- Trapping Diagnosis
33.98
327.40
657.47
649.11
1633.98
Table 2: Cost–performance tradeoff on RCA-100. We report joint accuracy (%) and mean token consumption per episode in thousands (K), averaged over all 103 cases. Token counts include input and output tokens for each component.
Figure 4: Success rates on LOCA-Bench as context grows from 8K to 256K across backbones.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Domain / Notes
Environment and interaction
st
Latent world state at step t
Not directly observable by the agent.
ht
Interaction history at step t
Contains previous actions and observations.
at,ot+1
Agent action and resulting observation
The environment returns ot+1 after action at .
T
Maximum interaction budget
Defined separately for each benchmark.
Belief representation and update
Appendix
Table 3: Summary of the principal notation used in PoS .
Boundary case
(pt,St,Rt)
Htgeo
Htmax
Ht
Stagnation without recurrence
(1,1,0)
1
0
0
Recurrence despite positive progress labels
(1,0,1)
1
0
0
Progress without gap closure
(1,0,0)
1
0
1
Appendix
Table 4: Health scores under three illustrative boundary configurations. Lower values indicate stronger trapping signals.
Benchmark
Evaluation Split
Task Category
# Cases
ALFWorld
valid_unseen
Place
41
Transform
75
Inspect
18
Total
134
LOCA-Bench
Official benchmark
Cross-source Workflows
175
Structured Analytics
175
Appendix
Table 5: Dataset statistics and evaluation splits.
Benchmark
Reported Metrics
Maximum Budget
Terminal Evaluation
ALFWorld
Category and Overall task success rates
50 text actions
The environment terminates the case when the goal is satisfied or another terminal state is reached. Success requires the environment’s goal-completion signal.
LOCA-Bench
Category and Overall task success rates
100 MCP tool calls
Calling claim_done triggers the official evaluator. A terminal score of at least 1.0 is counted as success.
RCA-100
FT, RCL, and Joint accuracy
50 decision turns; at most 150 tool calls
A valid submit_root_cause call terminates the case. FT and RCL are scored separately, and Joint requires both to be correct.
ClinDiag
Diagnosis accuracy
30 tool calls
A valid submit_diagnosis call terminates the case, after which the submitted diagnosis is evaluated against the reference diagnosis.
Appendix
Table 6: Evaluation metrics, interaction budgets, and termination conditions. All methods use the same budget within each benchmark.
Method
Decision Context
Implementation Source
Material Adaptation
Raw Trajectory
The task specification, current observation, and complete chronological action–observation trajectory.
Shared ReAct implementation.
No compression, retrieval, or auxiliary context-management operation.
ACON
A cumulative compressed summary together with the three most recent interaction turns in their original form; before compression is triggered, the full trajectory is retained.
Official implementation (commit d63f9ae18959 ) [ Kang et al., 2025 ] .
Runtime compression is preserved, while failure-driven guideline optimization and compressor distillation are omitted.
PACE
A chronological sequence of historical chunks represented at full, detailed, brief, or placeholder granularity according to predicted relevance, with the two most recent chunks retained in full and selected chunks recoverable through glimpse.
Reimplemented from the published method and Algorithm 1 [ Wei et al., 2026 ] .
The published repository was inaccessible during implementation; text-embedding-v4 replaces the original BGE-M3 encoder for relevance scoring.
HiAgent
Summaries of completed subgoals followed by the full action–observation trajectory of the active subgoal; original trajectories of completed subgoals remain retrievable when needed.
Official implementation (commit cebdd8e4eace ) [ Hu et al., 2025 ] .
The hierarchical working-memory mechanism is ported from AgentBoard to the shared benchmark interface, with benchmark-level summary and folding parameters.
LongHorizon-Harness
An audited task state containing requirements, artifacts, and facts, together with a bounded subtask contract for the current execution round.
Official implementation [ Ma et al., 2026 ] .
The native Manage–Execute–Audit architecture is preserved, with its benchmark-facing tool interfaces adapted to our common evaluation setup.
Appendix
Table 7: Decision contexts and implementation provenance of the evaluated baselines. All methods share the same benchmark interface and environment-action budget. The main differences lie in how interaction evidence is represented and exposed for subsequent decision making.
Parameter
Value
Belief update interval
1
Local audit interval
1
Global audit interval (LOCA-Bench)
32
Global audit interval (other benchmarks)
8
Recent transitions supplied for local auditing
4
Detection window size K
8
Appendix
Table 8: Hyperparameter configuration of PoS . Intervals are measured in interaction steps.
Table 9: Main results across four benchmarks (%). We report task success for ALFWorld and LOCA-Bench, failure triage (FT), root-cause localization (RCL), and joint accuracy for RCA-100, and diagnosis accuracy for ClinDiag. Env. Ops denotes Environment Operations.
Memory-augmented LLM agents tackle complex long-horizon tasks by recursively summarizing interaction trajectories into compact memory. However, existing approaches typically train these memory policies using outcome-based reinforcement learning, failing to localize where intermediate memory quality degrades. As interactions unfold, ambiguous recursive summaries progressively discard task-relevant information and introduce semantic noise. This exacerbates belief deviation, obscuring the agent's estimate of the latent task state and ultimately derailing long-horizon reasoning. We therefore argue that memory optimization should focus not merely on trajectory-level success, but on the clarity of the belief induced by intermediate summaries. To this end, we introduce Belief Entropy, a self-supervised proxy that probes how uncertain the model remains about the latent task state given its current memory. Based on this proxy, we propose Metacognitive Memory Policy Optimization (MMPO). Instead of relying only on sparse outcome-based signals, MMPO provides fine-grained, memory-specific supervision via explicitly penalizing summaries that induce high epistemic uncertainty. Experiments show that MMPO consistently outperforms existing methods on diverse long-horizon tasks, maintaining 97.1% performance even when scaled to 1.75M-token contexts.
Ziyan Liu, Zhezheng Hao, Yeqiu Chen +7
University of Science and Technology of China · Zhejiang University · Tencent
Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents. This gap is reflected in their limited performance on continual learning benchmarks such as ARC-AGI-3, especially when models are evaluated out of the box. Various agent harnesses have been proposed to close this gap, and each commits to a strategy for handling long sequences of observations, i.e., what information to save from the environment and how to load it into model context, a choice we argue is particularly consequential. Existing methods for context management face a significant tradeoff, as preserving more information makes retrieving relevant details less tractable. We propose PRO-LONG, a minimal context management framework built around programmatic memory for LLM agents in long-horizon, exploratory settings. PRO-LONG addresses the tradeoff by keeping a complete, structured interaction log and capitalizing on recent progress in coding agents to search this history efficiently. On the full ARC-AGI-3 public game set, PRO-LONG improves over a base coding agent by an average of 18.0 percentage points across frontier models, and matches or exceeds state-of-the-art specialized harnesses (up to 76.1% pass@1) while using 4.2-5.8x fewer tokens. With Fable 5, PRO-LONG achieves 97.4% best@2 at a total cost of $1,750. Relevant code and logs are available at https://github.com/alexisfox7/PRO-LONG.
Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses. To address these challenges, we propose FlowState, which treats execution state as memory that can be retained and revisited across requests, unifying current decision-making with the reuse of historical information. FlowState preserves semantically typed state nodes, their relations, and references to raw tool observations, separating persistent retention from on-demand access. Within a single execution loop, Incremental State Update (ISU) maintains the current state based on new inputs and feedback, while Progressive State Access (PSA) progressively reveals historical states and supporting evidence as needed during reasoning. Together, these mechanisms enable agents to reassess prior decisions in light of new information and guide subsequent actions. Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on τ3-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption by 43.2% and 40.6%. These results demonstrate the performance and efficiency advantages of FlowState on long-horizon tasks.