Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world. We introduce PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent's decision context. Each belief combines an estimate of the current world state with unresolved task requirements, making explicit what the agent still needs to learn and accomplish. To keep this belief reliable and actionable, PoS validates its consistency and monitors task progress to detect Belief Trapping, where the agent continues to act without making meaningful progress toward the goal. Recovery is then tailored to both the trapping pattern and the type of unresolved task requirement. Experiments on four benchmarks spanning execution and diagnosis show that PoS achieves the highest overall performance on every benchmark with all three LLM backbones. Ablations demonstrate the importance of consistency validation and recovery, while context-scaling experiments show resilience to context growth. Together, these results support belief construction and continual maintenance as a foundation for long-horizon context management beyond history retention and compression.
Figures & tables
Figure 1: Performance and motivation of PoS . Left: PoS outperforms the strongest baseline for each backbone across four benchmarks. Right: a real ALFWorld case contrasts implicit belief reconstruction that leads to Belief Trapping with explicit belief maintenance in PoS , which enables consistency validation and recovery toward task completion.
Figure 2: Overview of PoS . Belief Modeling constructs and incrementally updates a task-conditioned Belief from interaction evidence (Section 3.1 ). Belief-Guided Interaction conditions each action on the current Belief and an Active Gap, while the Belief Sentinel validates consistency and records task-relevant progress (Section 3.2 ). Trapping-Aware Recovery estimates Belief health, diagnoses heterogeneous trapping conditions, and composes recovery constraints to restore progress (Section 3.3 ).
Table 1: Main results across four benchmarks. We report overall task success for ALFWorld (ALF) and LOCA-Bench (LOCA), joint accuracy for RCA-100 (RCA), and overall diagnosis accuracy for ClinDiag (Clin), all in %. Best performance is bold. Fine-grained results in Table 9
Figure 3: Belief Trapping and recovery across benchmarks. Left: Percentage of episodes with detected trapping under each backbone. Middle: Distribution of trapping patterns by benchmark. Right: Final task performance over all evaluated cases with Qwen3.7-Plus, comparing PoS ’s factorized recovery with a generic recovery prompt.
Backbone
Method
Joint (%)
Tokens per episode (K)
Task Agent
Belief
Sentinel
Total
Qwen3.7-Plus
Raw Trajectory
24.27
355.70
0.00
0.00
355.70
PoS (Ours)
38.83
281.20
696.69
822.24
1800.13
- Consistency Validation
31.07
299.82
733.41
125.30
1158.53
- Trapping Diagnosis
33.98
327.40
657.47
649.11
1633.98
Table 2: Cost–performance tradeoff on RCA-100. We report joint accuracy (%) and mean token consumption per episode in thousands (K), averaged over all 103 cases. Token counts include input and output tokens for each component.
Figure 4: Success rates on LOCA-Bench as context grows from 8K to 256K across backbones.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Domain / Notes
Environment and interaction
st
Latent world state at step t
Not directly observable by the agent.
ht
Interaction history at step t
Contains previous actions and observations.
at,ot+1
Agent action and resulting observation
The environment returns ot+1 after action at .
T
Maximum interaction budget
Defined separately for each benchmark.
Belief representation and update
Appendix
Table 3: Summary of the principal notation used in PoS .
Boundary case
(pt,St,Rt)
Htgeo
Htmax
Ht
Stagnation without recurrence
(1,1,0)
1
0
0
Recurrence despite positive progress labels
(1,0,1)
1
0
0
Progress without gap closure
(1,0,0)
1
0
1
Appendix
Table 4: Health scores under three illustrative boundary configurations. Lower values indicate stronger trapping signals.
Benchmark
Evaluation Split
Task Category
# Cases
ALFWorld
valid_unseen
Place
41
Transform
75
Inspect
18
Total
134
LOCA-Bench
Official benchmark
Cross-source Workflows
175
Structured Analytics
175
Appendix
Table 5: Dataset statistics and evaluation splits.
Benchmark
Reported Metrics
Maximum Budget
Terminal Evaluation
ALFWorld
Category and Overall task success rates
50 text actions
The environment terminates the case when the goal is satisfied or another terminal state is reached. Success requires the environment’s goal-completion signal.
LOCA-Bench
Category and Overall task success rates
100 MCP tool calls
Calling claim_done triggers the official evaluator. A terminal score of at least 1.0 is counted as success.
RCA-100
FT, RCL, and Joint accuracy
50 decision turns; at most 150 tool calls
A valid submit_root_cause call terminates the case. FT and RCL are scored separately, and Joint requires both to be correct.
ClinDiag
Diagnosis accuracy
30 tool calls
A valid submit_diagnosis call terminates the case, after which the submitted diagnosis is evaluated against the reference diagnosis.
Appendix
Table 6: Evaluation metrics, interaction budgets, and termination conditions. All methods use the same budget within each benchmark.
Method
Decision Context
Implementation Source
Material Adaptation
Raw Trajectory
The task specification, current observation, and complete chronological action–observation trajectory.
Shared ReAct implementation.
No compression, retrieval, or auxiliary context-management operation.
ACON
A cumulative compressed summary together with the three most recent interaction turns in their original form; before compression is triggered, the full trajectory is retained.
Official implementation (commit d63f9ae18959 ) [ Kang et al., 2025 ] .
Runtime compression is preserved, while failure-driven guideline optimization and compressor distillation are omitted.
PACE
A chronological sequence of historical chunks represented at full, detailed, brief, or placeholder granularity according to predicted relevance, with the two most recent chunks retained in full and selected chunks recoverable through glimpse.
Reimplemented from the published method and Algorithm 1 [ Wei et al., 2026 ] .
The published repository was inaccessible during implementation; text-embedding-v4 replaces the original BGE-M3 encoder for relevance scoring.
HiAgent
Summaries of completed subgoals followed by the full action–observation trajectory of the active subgoal; original trajectories of completed subgoals remain retrievable when needed.
Official implementation (commit cebdd8e4eace ) [ Hu et al., 2025 ] .
The hierarchical working-memory mechanism is ported from AgentBoard to the shared benchmark interface, with benchmark-level summary and folding parameters.
LongHorizon-Harness
An audited task state containing requirements, artifacts, and facts, together with a bounded subtask contract for the current execution round.
Official implementation [ Ma et al., 2026 ] .
The native Manage–Execute–Audit architecture is preserved, with its benchmark-facing tool interfaces adapted to our common evaluation setup.
Appendix
Table 7: Decision contexts and implementation provenance of the evaluated baselines. All methods share the same benchmark interface and environment-action budget. The main differences lie in how interaction evidence is represented and exposed for subsequent decision making.
Parameter
Value
Belief update interval
1
Local audit interval
1
Global audit interval (LOCA-Bench)
32
Global audit interval (other benchmarks)
8
Recent transitions supplied for local auditing
4
Detection window size K
8
Appendix
Table 8: Hyperparameter configuration of PoS . Intervals are measured in interaction steps.
Table 9: Main results across four benchmarks (%). We report task success for ALFWorld and LOCA-Bench, failure triage (FT), root-cause localization (RCL), and joint accuracy for RCA-100, and diagnosis accuracy for ClinDiag. Env. Ops denotes Environment Operations.