Tackling complex real-world tasks can exceed the capabilities of a single large language model (LLM), motivating the use of multi-agent workflows that coordinate specialized agents to work together on these tasks. Recent methods train LLMs to construct better workflows from execution outcomes, but they optimize only the workflow generator, while the other agents that build or execute each workflow remain fixed even though every outcome depends on all of them. However, extending training beyond the generator is challenging: the agents are coupled, and a workflow's outcome is a single sparse score that cannot tell which agent causes a failure. We propose FloWright, which leverages the workflow as a harness to optimize workflows. By introducing a hierarchical, structure-aware reward paradigm, FloWright enables one role to self-evolve and two or more roles to co-evolve, with no additional models, labels, or executions. Considering the limitation that workflows are commonly trained and evaluated on data that a single agent can already handle, we further propose DataWright, an adaptive data hardening approach that converts existing datasets into workflow-level tasks with increased difficulty. Across document, slide, chart, code, math, and finance tasks, small open models trained with FloWright achieve improved performance by up to +7.41%, with co-evolving (+5.03%) more roles gaining more than optimizing one of them alone (+2.83%). Our project page: https://xhguo7.github.io/FloWright/.
Figures & tables
Figure 1: Upstream-Downstream Framework. Upstream, workflow generation can take on different structures, from Generator with its own Inventor skill to Generator working with other agents, drawing on pools of different granularity. Downstream, workflows take different structures and topologies to better solve the task.
Figure 2: FloWright for reinforcement learning via self-evolving harness. πρ is the policy of a role ρ∈R . For the j -th of m rollouts of a role being optimized, Gj∈G is the workflow that rollout is graded on, and the m rollouts span M distinct workflows (Tab. 7 ). Workflow generation and execution form a coupled on-policy learning loop: upstream proposes workflows, downstream executes them, and the verifiable harness turns execution outcomes into learning signals for the policies being optimized. Policy updates produce the next generation of workflows and executions, enabling self-evolution by harnessing the system’s own experience without additional supervision.
Method
In-Dist.
Out-of-Distribution
Out-of-Domain
Overall
ℓ=3
ℓ=3
ℓ=5
Mean
ℓ=3
ℓ=5
Mean
Baseline: Single Agent ( without tool or skill )
Qwen3.5-4B
10.87
11.39
6.12
8.15
7.89
5.12
6.56
7.42
Qwen3.5-9B
16.17
14.45
9.35
11.31
13.55
9.53
11.61
11.94
Baseline: Single Agent ( with tools and skills )
Qwen3.5-4B
21.13
20.75
15.12
17.28
25.02
19.35
22.29
20.70
Table 1: Evaluation on Four Evolution Modes. Mean accuracy (%) over the hardening strategies and levels each dataset supports (§ D.1 ). FloWright trains on the paired strategy at ℓ=3 of four datasets, which defines three regimes (§ D.5 ): in-distribution is those four training arms; out-of-distribution is the same four datasets under a strategy, a level, or both that training never sees (13 arms); and out-of-domain is eight datasets FloWright never trains on (27 arms). In-distribution is ℓ=3 by construction, since ℓ=3 is the training level and ℓ=5 is held out. The green shading is proportional to the gain over untrained backbones. 4B and 9B abbreviate Qwen3.5-4B and Qwen3.5-9B . (a)-(d) denote the four evolution modes defined in § 4 .
Figure 3: Test-Time Optimization with Reusable Prior. The leftmost bar of each group is the untrained Qwen3.5-4B workflow baseline, and every other bar generates workflows by distilling a reusable prior (§ 3.5 ). Internal and external distillation draw that prior from own experiences and from a stronger teacher (GPT-5.4) respectively. Meta distillation additionally distills successful experiences into few-shot cases, from which roles self-evolve via meta learning (1-shot and 5-shot). The rightmost bar composes train-time evolution with the 5-shot prior.
Figure 4: Algorithm Generalizability. Each algorithm optimizes Generator on Qwen3.5-4B .
Figure 5: Reward Ladder. Each layer adds one term (§ 3.3 ): format and answer alone, plus validity (§ E.3 ), plus the structure-aware credit (§ 3.2 ).
Figure 6: Workflow Topology. Generator writes workflows in schema, state machine, graph, or code (§ E.4 ). Both use the Qwen3.5-4B backbone.
Figure 7: Pool Granularity. Backbones are compared at two pool granularities (§ 3.1 ), and Trained 4B carries RL-trained Generator .
Figure 8: Pool Dynamics. A static pool stays fixed, a dynamic pool expands on demand, and a grown pool also carries components distilled from past experiences (§ 3.1 ).
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Preliminary Study. (a) Among the 121 datasets that the 20 workflow generation methods evaluate on: 85 of the 121 are single-question math, short-answer QA, or function-level coding, each of which a single competent agent can answer in one pass. (b) Comparing multi-agent workflows against a single agent given the same compute budget: workflows underperform the single agent on the original datasets, yet gain ↑39.28 on DataWright -hardened dataset.
Figure 10: LLM Judge Prompts. Each judge prompt carries its system instruction and its user message (§ D.4 ), trimmed with [...] . Scores are graded to 0 only when code execution failed (Eq. 12 ).
Task Type
Evaluation Metric
Definition
Question Answering
multiple choices, one answer
the chosen option equals the reference option
Eq. 7
multiple choices, multiple answers
the share of reference options named, lowered by every extra one
Eq. 7
list of items
the share of reference items named, lowered by every extra one
Eq. 7
numeric
equality at the reference’s precision, after normalizing format and units
Eq. 8
open-ended
exact match after normalization, otherwise the token-level F1 with the reference
Eq. 9
Appendix
Table 2: Evaluation Metrics. We summarize the evaluation metrics of different types of tasks covered by DataWright (§ D.4 ).
Dataset
Strategy
Task Type
Input
Test Questions
Training Questions
Document Understanding
MMLongBench-Doc ( Ma et al., 2024 )
Shared
mixed
file
960 / 820
–
Paired
mixed
file
1,089 / 1,090
–
Decoy
mixed
file
1,091 / 1,091
–
LongDocURL ( Deng et al., 2025 )
Shared
mixed
file
456 / 405
1,443 / 1,160
Paired
mixed
file
540 / 540
1,782 / 1,780
Appendix
Table 3: Data Statistics. Every dataset DataWright hardens, the strategies it supports, the task type its answers take, the input its questions carry, and the number of original questions each hardened set holds at level ℓ=3 and ℓ=5 (written ℓ=3 / ℓ=5 ). Counts are questions rather than bundles, since a task carries ℓ of them under Shared and Paired : bundling regroups the same questions into fewer and harder tasks as the level rises, and under Decoy the level changes only how many inputs surround the one that answers the question. A level is kept only when its test side holds at least 100 original questions (§ D.2 ), which is why two Shared levels are absent. A dash "–" in the training column marks a dataset held out for evaluation only, whose questions all serve the test side.
Role
Format Reward ( f )
Validity Reward ( vρ )
Upstream: Building a Workflow
Planner
its plan parses
helpfulness : the graded gain its plan brings over the same flow without one
Generator
its workflow parses
structure : the workflow is a valid graph over the pool(s)
Inventor ( Generator Skill )
every creation decision and authored component parses
deciding and authoring : reused components resolve to the pool, and newly authored ones instantiate and run
Downstream: Executing a Workflow
Downstream agent
its output parses
step format and step liveness : every node it runs reaches an answer of its own
Appendix
Table 4: Per-Role Validity. What each role’s own output needs to satisfy, split into the upstream roles that build a workflow and the downstream agent that executes it (Fig. 1 ). Every reward is deterministic, so no learned judge enters the reward (§ E.2 ).
Figure 11: Workflow Topologies. We show an example of the same hardened task represented in four topologies: a step schema, an agent graph, a state machine, and a block code. Each topology example is directly extracted from our evaluation outputs and trimmed with [...] . The four topologies differ only in where the control flow is written down, e.g. , a separate control_flow block, a typed edge, a transition, a block type, etc . Each is then converted into a shared directed graph structure, so execution, credit, and evaluation stay uniform without knowing which topology produced it (§ E.4 ).
Role
Format ( wf )
Validity ( wv )
Answer ( wa )
Credit ( wc )
Planner
0.1
helpfulness 0.1
0.8
0.1
Generator
0.1
structure 0.1
0.8
0.1
Inventor ( Generator Skill )
0.1
grounding 0.1 + authoring 0.1
0.7
0.1
Critic
0.1
helpfulness 0.3
0.6
0.1
Downstream agent
0.1
step format 0.025 + step liveness 0.075
0.8
0.1
Appendix
Table 5: Per-Role Reward Configuration. Every role instantiates the same hierarchical ladder (Eq. 3 ). The validity term is role-specific, and the credit penalty rides on top of the weighted sum (§ E.1 ).
Setting
Value
Note
Training Configuration
Policy optimization
GRPO / DAPO / CISPO
ablated; all three instantiate Eq. 4
Learning rate
1×10−6
constant, uniform for all settings
Group size m
8
rollouts per task
Batch size
8
tasks per step
Epochs per data pass
1
constant, uniform for all settings
Appendix
Table 6: Experiment Configuration. We summarize the experiment configurations for training and evaluation (§ 4 ). The same configuration applies across all four evolution modes and roles.
Symbol
Meaning
Introduced in
Tasks and Workflows
x∼X , xi
a task drawn from the task distribution; xi is the i -th task
§ 2 , 3
G=(V,E)
a workflow: a directed graph whose nodes V are agents, tools, or skills, and whose edges E carry the data and control flow
§ 2
G
the space of workflows
§ 3
y
the answer a workflow returns
Eq. 1
Harness and Signal
Appendix
Table 7: Notation. We summarize notations used in our paper, grouped by the concept they denote, from tasks and workflows (§ 2 ) to rollouts, credit, reward, optimization (§ 3 ), and data hardening (§ D ). In § 3 , the indices i , j , and k run over tasks, rollouts, and distinct workflows, respectively, and symbols indexed by ρ are specific to role ρ , e.g. , cjρ is the credit of role ρ in the j -th rollout.
Symbol
Meaning
Introduced in
Data Hardening
B , (q,z,y⋆)
a source dataset, and one of its samples: a question q , its input z , and its reference answer y⋆
Alg. 1
σ
a hardening strategy: Shared , Paired , or Decoy
§ D.1
ℓ
the hardening level: questions per task for Shared and Paired , and inputs per task for Decoy
§ D.1
Tσtrain , Tσtest
the hardened tasks that strategy σ builds from a dataset’s training split, and those it builds from its test split
Alg. 1
∣x∣
the number of sub-questions a hardened task x holds, one per original sample bundled into it
Alg. 1
Appendix
Table 19
Figure 12: Per-Arm Results Behind the Mean Scores. Complementing the mean accuracy scores in Tab. 1 , we show the accuracy (%) of each method on the 44 evaluation arms, grouped into the three regimes (Tab. 1 , § D.5 ): in-distribution , out-of-distribution and out-of-domain . We denote the hardening strategy and level as: S/P/D = Shared/Paired/Decoy, 3/5 = level. Color intensity encodes accuracy on a shared scale, and hue identifies the regime.
Figure 13: Roles Powered by Different Backbones. Accuracy (%) across three evaluation regimes over all 44 arms (§ 4 ). Single Backbone uses one backbone for every role. Untrained Mix uses a 4B or 9B backbone to specific roles, with no role optimized. Trained Role ( ∗ ) uses the RL-trained 4B role, with Qwen3.5-9B powering other roles. Gen and Skill name the Generator and the skill it co-evolves with, and others covers every remaining role.
Figure 14: Evolution of Different Roles. Accuracy (%) across three evaluation regimes over all 44 arms (§ 4 ). Every role runs Qwen3.5-4B , with the only variable as which role the harness optimizes. Single-Role Self-Evolution optimizes the Generator , skill, and downstream agent individually. Co-Evolution optimizes two or more together (§ 3.4 ).
Figure 15: Meta Distillation Across Backbones. Accuracy (%) across three evaluation regimes over all 44 arms (§ 4 ). Each backbone appears as an untrained Baseline , compared with +Distill providing the meta prior distilled at five shots (§ 3.5 ).
Figure 16: Execution Reward by Task Type. Δacc (%) against untrained baselines, on the four coding arms (TACO and DS-1000) and over all 44 arms (§ 4 ). Every other ladder weight stays fixed (Tab. 5 ).