Organizations: The Chinese University of Hong Kong, Shenzhen, China · Dalian University of Technology, China · Fudan University, China · University of Oxford, UK · North China Electric Power University, China · The University of Texas Health Science Center at Houston, USA
In recent years, LLM-based multi-agent systems have been widely applied to orchestrate tool-using agents into executable communication graphs. However, existing self-evolving orchestration still faces key challenges, including post-hoc evolution that revises the team only after the trajectory ends, credit diffusion that gives every action the same terminal advantage under confounded baselines, and skill admission that is uncalibrated and never retired. To address these challenges, we propose EvoSteer, a new paradigm of Online Self-Evolving Graph Orchestration -- the orchestrator builds a running team and repairs its plausible but failing steps from execution features and a learned value estimate. To support this paradigm, we introduce Anchored Trajectory Balance (AnchorTB), a regression-style flow-matching loss that assigns each orchestration action a coefficient by balancing subtrajectories against a frozen reference. Built on the learned flow, we further propose Validated Skill Admission, in which a candidate skill is tried before promotion and promoted only if paired evidence passes a sequential test under a shared nominal testing budget. Moreover, AnchorTB combines measured task-level reference reward statistics with prefix-dependent corrections. Experimental results on twelve datasets show that EvoSteer significantly outperforms baselines across question answering, mathematical reasoning, code generation, and interactive decision making. Our code is available at https://github.com/beita6969/evosteer.
Figures & tables
Figure 1: Outputs that look plausible can still be headed for failure. EvoSteer reads a low continuation estimate off the execution record and edits the running team, here routing the checker’s report back to the solver; dashed: the same team left unedited.
Figure 2: Three lines of work and ours. (a) Post-hoc evolution revises the team only after the run. (b) Outcome-driven optimization spreads one terminal reward over every step. (c) Skill evolution admits by judges or replay. (d) EvoSteer edits the team inside the same run, deciding from measured execution features and a reference value estimate, and admits a skill only after a sequential test.
Figure 3: EvoSteer architecture. Top: πθ builds and runs the team from execution features ft and reference estimate v^k ; rerun and new edges are ordinary actions. Bottom left: AnchorTB balances subtrajectories against ρ ; u~,δ abbreviate trajectory-specific flow estimates and residuals. Bottom right: paired rollouts and a sign test under a shared nominal α budget promote or retire candidate σ .
Figure 4: AnchorTB on a three-action history: ① score each action against the frozen reference; ② anchor each state with a measured, stop-gradient anchor (blue) and a learned residual (pink); ③ form all K3=6 span residuals; ④ sum each action’s spans into its coefficient.
Variant
IID
OOD
HotpotQA
NQ-Open
MedQA
AIME
MBPP+
ALFWorld
TriviaQA
MuSiQue
GPQA
MATH
SWE
WebShop
Ans EM
Ans EM
Acc.
Acc.
Pass@1
SR
Ans EM
Ans EM
Acc.
Acc.
Resolved
SR
Qwen3.5-9B (frozen)
57.66
23.59
71.41
48.67
78.91
46.88
44.22
39.22
61.72
89.06
15.94
32.66
EvoSteer arch., untrained
82.34
73.28
82.97
53.33
87.66
73.59
90.47
71.88
75.31
90.00
29.53
69.84
Fixed orchestration paradigms
Single agent with tools
75.94
67.19
80.78
48.67
85.63
70.16
86.25
61.41
70.31
89.84
27.03
66.09
Table 2: Component ablation and paradigm comparison (untrained: initial πθ ). Paradigms (same executor and budget) fix the team before execution: a ReAct-style agent with all tools, a hand-designed template, a planner that writes the graph once, and a workflow searched offline on training tasks. Ablations: − interleaved execution builds the graph first; − execution features drops f from the state of πθ ; − learnable repair masks rerun and drop ; − reference value head withholds v^k and Δv^ ; − measured flows learns uq as a free scalar; − flow corrections turns off c(g) and bψ ; − skill evolution runs without any skills; − sequential validation admits after k=5 successes.
Figure 5: Backbone transfer and training dynamics. (a) IID scores of six other frozen executors (dashed) and with the same trained orchestrator (solid); the radius is linear over 20–80 on the inner half and 80–100 on the outer half. (b) Accuracy of πθ and of the frozen ρ , and loss against the anchor-only level, over 240 steps. (c) OOD scores per task domain.
Figure 6: Objective comparison and mechanism analysis (definitions in Appendix D ). (a) OOD scores and cost per training objective. (b) Row minus column, IID mean (pp); TTB: tempered TB. (c) AUC for predicting reference-rollout correctness. (d) Paired skill gain, admitting all or only validated candidates. (e) In-run edits by type and outcome. (f) Gain of value-guided replanning in training.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Group
Setting
Backbones
orchestrator, reference, executor: one shared base model ( Qwen3.5-9B ); skill author: DeepSeek-V4-Flash, called only in author windows
Adapter
LoRA rank 64, α=128 , on q/k/v/o_proj, out_proj, and in_proj_qkv; AdamW ( β1=0.95 ), lr 5×10−6
Heads
bψ : zero-initialized MLP on [hρ;f] ; value head: two-layer MLP on 30+6 inputs (one task-type slot per IID benchmark)
AnchorTB
β=2 ; all subtrajectories equally weighted; gradient-norm clip 1.0
at most three validated skills and one candidate per task type
Appendix
Table 3: Configuration of the frozen implementation used for the experiments.
IID avg.
OOD avg.
Skill author
Ans EM
Acc.
Ans EM
Acc.
DeepSeek-V4-Flash (main setting, Table 1 )
88.67
87.42
90.08
78.55
GPT-5.6-Luna
90.16
89.28
90.55
79.34
Qwen3.5-9B (the executor)
87.81
86.64
88.98
76.99
Appendix
Table 4: EvoSteer with different skill authors: five-run means averaged as in Table 1 (Ans EM over the question-answering benchmarks, Acc. over the others).
In recent years, a variety of powerful LLM-based agentic systems have been applied to automate complex tasks through task orchestration. However, existing orchestration methods still face key challenges, including strategy collapse under reward maximization, high gradient variance with opaque credit assignment, and unguided skill evolution whose decisions are typically made by directly prompting an LLM to judge rather than derived from principled training signals. To address these challenges, we propose SkillFlow, a flow-based framework that takes a trainable Supervisor as the agent and a structured environment with dynamic skill library and frozen executor, automating task orchestration through multi-turn interaction. SkillFlow employs Tempered Trajectory Balance (TTB), a regression-based flow-matching loss that samples trajectories proportional to reward, preserving diverse orchestration strategies rather than collapsing to a single mode. The same flow objective yields a jointly learned backward policy that provides transparent per-step credit assignment at zero additional inference cost. Building on these flow diagnostics, a recursive skill evolution mechanism determines when to evolve, what skills to create or prune, and where decision gaps lie -- closing the loop from training signal to autonomous capability growth. Experimental results on 14 datasets show that SkillFlow significantly outperforms baselines across question answering, mathematical reasoning, code generation, and real-world interactive decision making tasks. Our code is available at https://anonymous.4open.science/r/SkillFlow-E850.
Mingda Zhang, Tiesunlong Shen, Haoran Luo +4
The Chinese University of Hong Kong, Shenzhen · National University of Singapore · Nanyang Technological University +1
Large language models (LLMs) enable autonomous agents for reasoning, planning, and tool use. Recent systems increasingly organize these agents as graphs of specialized, interconnected nodes. Although graph-based orchestration supports flexible decomposition and coordination, it creates a key challenge: \textbf{attention allocation}. As workflows grow, existing approaches often execute graph components uniformly, wasting resources on irrelevant or low-impact tasks. We introduce \textbf{Attention Orchestration}, a paradigm that extends Transformer-style attention from token representations to workflow-level agent coordination. Our framework, \textbf{Adaptive Goal-aware Attention Orchestration (AGAO)}, dynamically estimates agent importance based on user objectives, graph dependencies, and computational constraints. AGAO combines three components: (1) goal-aware attention, measuring semantic relevance between user goals and agent capabilities; (2) topology-aware attention, modeling structural dependencies in agent graphs; and (3) resource-aware attention, allocating budgets and execution priorities across heterogeneous agents. Together, these mechanisms transform static agent graphs into adaptive systems that focus computation on goal-critical reasoning paths. Experiments across diverse multi-agent workloads show that AGAO improves task effectiveness while reducing unnecessary computation, latency, and token consumption compared with existing graph-based execution strategies. Our work establishes \textbf{Attention Engineering} as a direction for scalable, intelligent multi-agent systems. Code: https://github.com/MingzhouFan97/AGAO.
Large language model (LLM)-based agents have demonstrated strong capabilities in complex reasoning and problem solving through multi-step interactions, yet most deployed agents remain behaviorally static, with knowledge acquired during execution rarely translating into systematic improvement over time. In response, a growing line of work on self-evolving agents explores how agents can improve through experience during deployment, but most existing approaches either rely on ad hoc reflection limited to single-task correction or adopt unstructured memory that accumulates fragmented experience with delayed usability. To address this limitation, we introduce EXG, an experience graph framework for self-evolving agents that explicitly organizes accumulated successes and failures into a structured, relational representation. EXG is the first experience graph designed for self-evolving agents, supporting both online, real-time graph growth during execution for immediate cross-task experience reuse, and offline reuse of a consolidated experience graph as an external memory module. This design also enables EXG to serve as a plug-and-play component for existing self-evolving agents, organizing prior experience into a unified experience graph and improving both solution quality and resource efficiency as deployment progresses. Extensive experiments across code generation and reasoning benchmarks show that EXG attains more favorable performance-efficiency trade-offs than reflection- and memory-based baselines in both online and offline evaluations. Our results suggest that structuring experience as a graph provides a principled foundation for scalable and transferable self-evolving agent behavior.
Yuxin Jin, Siyuan Zhang, Hanchen Wang +3
University of Technology Sydney Sydney, Australia · The University of New South Wales Sydney, Australia