Organizations: School of Electronic Information, Wuhan University, Wuhan 430079, China · School of Remote Sensing and Information Engineering, Wuhan University, Wuhan 430079, China · State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430079, China
Large language model agents increasingly rely on compound programs for retrieval, tool use, reasoning, and verification, yet their failures often arise from local procedural decisions. Existing reinforcement-learning and prompt-optimization approaches typically rely on scalar rewards or repeatedly modify entire prompts, making it difficult to capture and reuse procedural improvements while preserving a frozen agent. To address this problem, We propose Natural-Language Policy Gradients (NLPG), an external policy-memory method for improving a fixed agent without changing its model parameters or program structure. NLPG diagnoses execution traces, propagates downstream feedback backward through the module graph, and converts recurring failures into route-local natural-language corrections that are aggregated into bounded policy updates for subsequent executions. Across six benchmarks covering memory, reasoning, instruction following, and evidence verification, NLPG also outperforms the strongest listed baseline for each benchmark by 8.71 percentage points on average. These results provide evidence that evaluated procedural experience can be transformed into local and interpretable policy updates, enabling continual improvement of frozen agents.
Figures & tables
Figure 1: Conceptual comparison of global prompt-level updates and NLPG’s route-indexed policy updates.
Figure 2: Overview of NLPG. (A) Inference and failure diagnosis. The compound agent retrieves module-local policies from the current policy memory Mt and injects them into the corresponding module calls. (B) Policy-memory update. NLPG aggregates routed language gradients across evaluated trajectories, scores and selects reusable instructions, and stores them under the corresponding program modules and task categories to construct Mt+1 .
LLM
Method
1. Multi-Hop
2. Temporal
3. Open-Domain
4. Single-Hop
Overall
GPT-4o-mini
MemoryOS
56.50
37.18
40.28
62.43
54.70
Mem0
58.16
55.45
40.62
66.71
61.00
MemU
62.41
33.96
46.88
72.77
61.15
MemOS
69.15
72.27
60.42
81.45
75.87
HiMem
70.92
74.77
54.86
89.22
80.71
Zep
71.99
74.45
66.67
88.11
81.06
Table 1: LoCoMo results under the aligned evaluation protocol. Parenthesized values show improvements over the strongest baseline in each column.
Qwen3-8B
Method
HotpotQA
IFBench
HoVer
3-task Avg.
Improvement
Baseline
42.33
36.90
35.33
38.19
–
GRPO
43.33
35.88
38.67
39.29
+1.11
MIPROv2
55.33
36.22
47.33
46.29
+8.11
GEPA
62.33
38.61
52.33
51.09
+12.90
GEPA+Merge
64.33
28.23
51.67
48.08
+9.89
Table 2: Cross-model generalization under the GEPA evaluation suite Agrawal et al. (2026) . NLPG is evaluated with Qwen3-8B Yang and others (2025) and GPT-4.1 Mini OpenAI (2025) on HotpotQA, IFBench, and HoVer.
Figure 3: Cross-benchmark accuracy comparison. Panels (a) and (b) show transfer results for Qwen3-8B and GPT-4.1 Mini, respectively, while panel (c) reports category-wise LoCoMo accuracy. Consistent with Tables 2 and 1 , NLPG achieves the best performance on all transfer tasks for both backbones and improves every LoCoMo category, with the largest gains on multi-hop, temporal, and open-domain questions. The strongest improvements occur on procedurally compositional questions, suggesting that route-specific policy updates improve decisions about query preservation, memory combination, and retrieval expansion rather than simply increasing context size.
Figure 4: IFBench ablations by instruction category. (a) Strict accuracy (%). (b) Difference from full NLPG in percentage points. Hatched cells contain fewer than 10 examples.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Input → Output
Role in NLPG
Policy loading
Mt→ route-indexed policy candidates
Load the active procedural policy version.
Policy selection
Task input and route → top- K instructions
Deduplicate, rank, and select the policy context for the invocation.
Fixed execution
Task input and selected policy → execution trace
Run the original benchmark program without changing its structure.
Evaluation
Execution trace →zi
Apply the benchmark’s released checker or judge.
Failure analysis
(τi,zi)→ failure feedback
Identify the deviation and the routes relevant to its correction.
Policy update
Failure feedback → textual corrections
Generate and consolidate reusable route-specific instructions.
Appendix
Table 3: Conceptual data flow of the NLPG execution and update process.
Benchmark
Route representation
Policy injection point
K
LoCoMo
Question category
Retrieval and answer guidance
4
HaluMem
Inferred question family
Memory retrieval and answer policy
4
LongMemEval
Question type with fallback family
Retrieval and answer guidance
4
HotpotQA
Module routes in the two-hop program
Summarization, second-hop query, and answer modules
3
IFBench
Instruction identifier and instruction family
Response-generation context
3
HoVer
Hop-specific summary and query modules
Multi-hop summary and query generation
3
Appendix
Table 4: Execution routes and policy injection points used by the evaluated benchmark programs.
Benchmark
Policy construction
Policy selection
Final test
Separation key
LoCoMo
Non-test conversation–user groups
Disjoint conversation–user groups
152 first-user questions from held-out groups
Conversation and user ID
HaluMem
Non-test user records
Disjoint validation user records
164 questions from held-out user records
User and record ID
LongMemEval
Non-test history–question groups
Disjoint validation groups
Held-out LongMemEval-S questions
History and question ID
HotpotQA
150 train examples
300 validation examples
300 test examples
Example ID
IFBench
Construction portion of calibration data
Disjoint selection portion
150 held-out examples
Prompt/example ID
HoVer
Construction portion of 150 train examples
Disjoint selection portion
300 test examples
Claim ID
Appendix
Table 5: Independent held-out evaluation protocol. Construction, selection, and final-test partitions are mutually disjoint under the indicated grouping key.
Figure 5: Paired answer-stage ablations on LoCoMo ( n=152 ). Points denote accuracy changes relative to the full model, and horizontal bars show paired-bootstrap 95% confidence intervals. Negative values favor the full model, while the vertical line denotes no effect. Reported p -values are from two-sided exact McNemar tests. Retrieval results are fixed and cached across all variants.
Figure 6: Category-level changes in accuracy, measured in percentage points relative to the full configuration, for answer-stage ablations on the held-out LoCoMo set with fixed cached retrieval. Negative values indicate that the ablated configuration performs worse than the full configuration. The numbers in parentheses denote the number of questions in each category. The Open Domain column is interpreted cautiously because this category contains only 13 questions.
Method
Memory Integrity
Memory Accuracy
QA Accuracy
MemoBase
14.55
92.24
35.53
Supermemory
41.53
90.32
54.07
Mem0
42.91
86.26
53.02
ProMem
73.80
89.47
62.26
NLPG
70.25
96.20
85.37
Appendix
Table 6: Results on HaluMem under the aligned benchmark evaluation protocol. The three metrics are reported separately because they measure different properties of memory construction and downstream use.
Method
QA Accuracy
NativeRAG
65.09
Mem0
43.31
LightMem
68.64
ProMem
69.57
NLPG
71.80
Appendix
Table 7: Results on LongMemEval under the aligned question-answering protocol.
Setting
Value
Dataset
HotpotQA fullwiki, GEPA-aligned split policy
Evaluation split
Test; 300 examples; subset seed 1
Data partition
40% test, 40% validation, 20% train
Retrieval
Two-hop retrieval; top- k=5 for each hop
NLPG policy context
At most 3 instructions per module; maximum 1,000 characters
Time filtering
Disabled
Appendix
Table 8: Main configuration of the 300-example HotpotQA evaluation.
Metric
Strict
Loose
Paired records
143
143
Pass before repair
52/143 (36.36%)
57/143 (39.86%)
Pass after repair
59/143 (41.26%)
64/143 (44.76%)
Repair success
8/91 (8.79%)
8/86 (9.30%)
Regression
1/52 (1.92%)
1/57 (1.75%)
Net gain
+7
+7
Appendix
Table 9: Official paired repair evaluation on the mixed IFBench calibration subset. Repair success is computed over responses that fail the corresponding checker before repair. Regression is computed over responses that pass the corresponding checker before repair.
Figure 7: Feedback quality and repair evaluation on IFBench. (a) Paired official evaluation, showing pass rates before and after repair under strict and loose checkers. (b) Gain and regression, showing repair-success and regression rates. (c) Strict repair success by official instruction family.
Observable
Without policy
With policy
Question
Trey Anastasio and Glenn Bidmead have which mutual occupation?
Supporting-document recall
1.0
1.0
Hop-2 query
What are the shared occupations of Trey Anastasio and Glenn Bidmead?
What occupation do Trey Anastasio and Glenn Bidmead share?
Predicted answer
musician
guitarist
Exact-match correctness
false
true
Appendix
Table 10: Successful controlled policy intervention. Retrieval coverage is unchanged; only the module-local policy differs between the paired runs.