Prompt optimization has become a practical way to improve the performance of Large Language Models (LLMs) without retraining. However, most existing frameworks treat evaluation as a black box, relying solely on outcome scores without explaining why prompts succeed or fail. Moreover, they involve repetitive trial-and-error refinements that remain implicit, offering limited interpretability or actionable guidance for systematic improvement. In this paper, we propose MA-SAPO: a new Multi-Agent Reasoning for Score Aware Prompt Optimization framework that links evaluation outcomes directly to targeted refinements. Specifically, in the Training Phase, multiple agents interpret evaluation scores, diagnose weaknesses, and generate concrete revision directives, which are stored as reusable reasoning assets. In the Test Phase, an analyzer agent retrieves relevant exemplars and assets for a new prompt, and a refiner agent applies evidence-based edits to improve the prompt and its response. By grounding optimization in structured reasoning, MA-SAPO ensures edits are interpretable, auditable, and controllable. Experiments on the HelpSteer1/2 benchmarks show that our framework consistently outperforms single-pass prompting, retrieval-augmented generation, and prior multi-agent methods across multiple evaluation metrics.
Figures & tables
Figure 1: MA-SAPO pipeline overview. Three training agents convert annotated prompt–response pairs into reusable Reasoning Assets Ri=(Ci,Di,Ei) . At test time, retrieved assets guide the Analyzer Gana and Refiner Gref to produce an optimized prompt p^ with higher scores across all five evaluation dimensions.
Dataset
Train
Val
Multi-turn
Preference Ann.
HelpSteer1
35.3k
1.79k
×
×
HelpSteer2
20.3k
1.04k
✓
✓
Table 1: Statistics of the HelpSteer1/2 datasets. Both share the same annotation schema of five quality dimensions scored on a 0–4 scale.
Methods
GPT-4o
Llama3-8B
Help
Corr
Coh
Comp
Verb
Avg
Help
Corr
Coh
Comp
Verb
Avg
HelpSteer1
Direct Generation
0.3216
0.3866
0.7583
0.2951
0.4215
0.4366
0.2549
0.3247
0.7054
0.2702
0.4062
0.3927
RAG (sparse, k=10 ) ( 2020 )
0.3751
0.4402
0.7871
0.3116
0.4779
0.4784
0.2930
0.3594
0.7452
0.3007
0.4377
0.4272
Chain-of-Thought (CoT) ( 2022 )
0.3223
0.3876
0.7595
0.2935
0.4192
0.4364
0.2359
0.3077
0.6906
0.2627
0.3967
0.3787
Role Assignment ( 2023 )
0.3988
0.4679
0.8024
0.3427
0.5008
0.5025
0.2887
0.3516
0.7287
0.3164
0.4517
0.4274
Table 2: Main results on HelpSteer1 and HelpSteer2 . Columns report the five HelpSteer metrics (Help, Corr, Coh, Comp, Verb) normalized to [0,1] and their mean ( Avg ). Methods are grouped into (i) single-pass prompting, (ii) retrieval-augmented generation (no reasoning assets), and (iii) multi-agent frameworks. For each dataset and backbone block, the best value in each metric column is shown in bold , and the second-best is underlined .
Methods
GPT-4o
Llama3-8B
Help
Corr
Coh
Comp
Verb
Avg
Help
Corr
Coh
Comp
Verb
Avg
HelpSteer1
k=1
0.5182
0.6247
0.8543
0.4943
0.7349
0.6453
0.4045
0.4769
0.8268
0.4729
0.8446
0.6051
k=2
0.5211
0.6287
0.8591
0.4960
0.7318
0.6473
0.4002
0.4747
0.8236
0.4726
0.8384
0.6019
k=4
0.5204
0.6266
0.8606
0.4965
0.7369
0.6482
0.4028
0.4778
0.8272
0.4708
0.8389
0.6035
Test Agents Combination
0.4672
0.5415
0.8619
0.3670
0.6403
0.5756
0.4649
0.5203
0.8323
0.4363
0.7249
0.5957
Table 3: Ablation results over agent number k and agent-combination testing on HelpSteer1 and HelpSteer2 . Best scores are shown in bold and second-best scores are underlined .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Attribute
Description
Helpfulness
Overall helpfulness of the response to the prompt.
Correctness
Inclusion of all pertinent facts without errors.
Coherence
Consistency and clarity of expression.
Complexity
Intellectual depth required to write the response.
Verbosity
Amount of detail included in the response relative to what is asked.
Appendix
Table 4: Overview of HelpSteer1/2 annotation dimensions. Each scored on a 0–4 scale.
Method
Usefulness
Accuracy
Consistency
Mean
Single-Agent
3.64
3.63
3.81
3.69
MA-SAPO (Ours)
3.89 ∗
3.87 ∗
4.02
3.93
Δ (Improvement)
+0.25
+0.24
+0.21
+0.24
Appendix
Table 5: Human evaluation of reasoning quality. MA-SAPO outperforms the single-agent baseline in usefulness, factual accuracy, and consistency. Scores are averaged on a 1–5 scale; asterisks denote significance levels from paired t -tests ( ∗p<0.05 , n=30 ).
Prompt template: Metric explainer agent
Role: You are an evaluation explainer. Task: • Explain WHY each of the five scores (0–4) was assigned to the model response. • Provide a cohesive paragraph that bridges the gap between numerical scores and actionable improvements. Inputs: • [USER_PROMPT] : {prompt_text} • [MODEL_RESPONSE] : {model_response} • [SCORES] : {5 metrics (helpfulness, correctness, coherence, complexity, verbosity)} Strict Rules: • Output must be a single paragraph in plain text. • Cover metrics in order: helpfulness → correctness → coherence → complexity → verbosity. • Ground every rationale specifically in the provided response evidence. • Maintain a neutral and specific tone.
Appendix
Table 6: Prompt template for the metric explainer agent.
Prompt template: Diagnostician agent
Role: You are a precise evaluation doctor for LLM outputs. Task: • Diagnose the primary strengths and weaknesses of the response based on metric explanations. • Prescribe concrete fixes to resolve identified reasoning or alignment gaps. Inputs: • [USER_PROMPT] : {prompt_text} • [MODEL_RESPONSE] : {model_response} • [SCORES] : {5 metrics (helpfulness, correctness, coherence, complexity, verbosity)} • [REASONING_PARAGRAPH] : {reasoning_paragraph} Strict Rules: • Rely strictly on the prompt, response, and provided reasoning. • Address each metric once with a brief, specific rationale. • Flag any mismatches between the numerical scores and the qualitative reasoning. • Conclude with a crisp prescription for the required response elements.
Appendix
Table 7: Prompt template for the diagnostician agent.
Prompt template: Action synthesizer agent
Role: You are a suggestion provider. Task: • Synthesize the diagnosis into a single actionable prompt suggestion. • Focus on enhancing the prompt’s structural and contextual depth without altering the original intent. Inputs: • [USER_PROMPT] : {prompt_text} • [MODEL_RESPONSE] : {model_response} • [SCORES] : {5 metrics (helpfulness, correctness, coherence, complexity, verbosity)} • [DIAGNOSIS] : {diagnosis_text} Strict Rules: • Output ONLY the improvement suggestion (no conversational filler). • Maintain the core topic and original objective of the user prompt. • Focus on actionable directions (e.g., adding constraints, requesting specific formats). • Balance the five metrics to guide the model toward a higher-quality response.
Appendix
Table 8: Prompt template for the action synthesizer agent.
Prompt template: Analyzer agent
Role: You are the analyzer agent. Task: • Conduct an independent evaluation of the prompt’s inherent qualities. • Integrate insights from the three previous agents (Metric Explainer, Diagnostician, Action Synthesizer) into a unified analysis. Inputs: • [USER_PROMPT] : {prompt} • [RETRIEVED_CONTENT] : {retrieved_content} (Aggregated agent outputs) Strict Rules: • Prioritize the Action Synthesizer’s output when synthesizing final improvement points. • Eliminate redundant observations across agent outputs. • Ensure all proposed improvements are concrete and actionable. • Output the final analysis in a structured, cohesive paragraph.
Appendix
Table 9: Prompt template for the analyzer agent.
Prompt template: Refiner agent
Role: You are the refiner agent. Task: • Optimize the original prompt based on the Analyzer Agent’s feedback. • Produce a final, high-performance prompt that maximizes clarity and task effectiveness. Inputs: • [USER_PROMPT] : {prompt} • [ANALYZER_FEEDBACK] : {analyzed_result} Strict Rules: • Directly incorporate the prioritized improvements from the feedback. • Preserve the original intent and task objective perfectly. • Remove any unnecessary elaboration to ensure the prompt remains concise. • Output ONLY the final optimized prompt.
Large language model (LLM)-based Multi-agent systems (MAS) have shown promise in tackling complex collaborative tasks, where agents are typically orchestrated via role-specific prompts. While the quality of these prompts is pivotal, jointly optimizing them across interacting agents remains a non-trivial challenge, primarily due to the misalignment between local agent objectives and holistic system goals. To address this, we introduce MASPO, a novel framework designed to automatically and iteratively refine prompts across the entire system. A core innovation of MASPO is its joint evaluation mechanism, which assesses prompts not merely by their local validity, but by their capacity to facilitate downstream success for successor agents. This effectively bridges the gap between local interactions and global outcomes without relying on ground-truth labels. Furthermore, MASPO employs a data-driven evolutionary beam search to efficiently navigate the high-dimensional prompt space. Extensive empirical evaluations across 6 diverse tasks demonstrate that MASPO consistently outperforms state-of-the-art prompt optimization methods, achieving an average accuracy improvement of 2.9. We release our code at https://github.com/wangzx1219/MASPO.
Zhexuan Wang, Xuebo Liu, Li Wang +4
Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen, China.
System prompt optimization improves agent behavior without modifying the underlying model, yielding human-readable, model-agnostic instructions. Existing methods build a prompt agent that refines task agents' system prompts, yet leave the prompt agent's own system prompt hand-engineered and fixed. We propose Self-Evolving Prompt Optimization (SePO), which treats the prompt agent's own system prompt as an optimization target alongside task agents' system prompts. SePO adopts a self-referential design. A single prompt agent improves both task agents' system prompts and its own under an open-ended evolutionary search that maintains an archive of candidate prompts as stepping stones. Training proceeds in two stages: pre-training evolves the prompt agent on a multi-task pool, and fine-tuning then applies it to a target task. Across five benchmarks spanning math (AIME'25), abstract reasoning (ARC-AGI-1), graduate-level science (GPQA), code generation (MBPP), and logic puzzles (Sudoku), SePO consistently outperforms Manual-CoT, TextGrad, and MetaSPO, improving the average accuracy by 4.49 points compared to Manual-CoT. The prompt optimization skill from pre-training also generalizes to tasks beyond the pre-training mixture, rather than memorizing per-task prompts.
Wangcheng Tao, Han Wu, Weng-Fai Wong
National University of Singapore · City University of Hong Kong
Prompt optimizers automate the search for prompts that improve language-model performance, but existing methods rely on a predefined optimization procedure: the algorithm determines which candidates to explore and how the search progresses, while the language model generates or refines prompt proposals. We introduce RLMOpt, a prompt optimizer that makes the search policy itself language-model-driven through a recursive language model (RLM). The RLM agent operates over a tool-based environment, inspecting task information, analyzing failures, generating candidates, allocating evaluation budget, and deciding when to stop. A deterministic harness complements the agent by enforcing objective scoring, Pareto-based selection, and regression constraints. We evaluate RLMOpt across four benchmarks spanning structured clinical information extraction (Chia), multi-hop question answering (HotpotQA), verifiable instruction following (IFBench-2025), and multi-turn tool-calling agents (BFCL). In a matched comparison at a single seed, RLMOpt obtains the best held-out score on all four benchmarks and leads the four-task mean (0.610 against 0.589 for GEPA). Repeating each benchmark across seeds yields 11 matched benchmark-seed comparisons, in which RLMOpt outperforms GEPA in 9 cases. Across all 11 runs, it never produced a prompt that underperformed its seed, whereas GEPA fell below its starting point twice. It is also more efficient, achieving these results with fewer search rollouts while producing prompts that are 27-79% the size of those produced by GEPA. Our results further show that optimization gains are determined primarily by the headroom available in the seed prompt, rather than by the search budget. Efficient optimization therefore depends on reaching the available headroom reliably and with minimal search