Large language models (LLMs) increasingly construct multi-agent workflows that decompose a complex task and assign specialist agents from a pool. However, building such a workflow well remains challenging: how finely to divide the task, which agent to trust with each subtask, and when to create a new specialist are all critical decisions a workflow constructor needs to settle up front. Thus, whether each subtask succeeds remains unknown until the workflow runs. Yet, improving a workflow is costly. Locating a fault usually requires a reference answer, a graded outcome, or a trained assessor, and the fix is applied to the whole workflow through re-execution, re-search, or retraining. We propose InFlowOp, which prices every decision in one label-free cost that weighs how well an agent's competence meets what a subtask demands against how much that agent takes to run. Before execution, InFlowOp bidirectionally determines the granularity of task decomposition and agent assignment following from the cost rather than from a fixed template. During execution, InFlowOp corrects a fault with the cheapest move via the same cost that serves the workflow both as it is built and as it runs. Facing the workflow-level evaluation challenge, we introduce Braid, a benchmark whose tasks require multi-agent coordination beyond single-agent capability. Across various domains and backbones, InFlowOp outperforms single agent baselines by up to +11.97%, achieving +9.64% with in-flow optimization. Our project page: https://xhguo7.github.io/InFlowOp/.
Figures & tables
Figure 1: InFlowOp for In-Flow Optimization. Complex tasks challenge single agents with long-context reasoning, while multi-agent systems introduce execution failures and coordination overhead. InFlowOp optimize agent behaviors in-flow, enabling efficient collaboration at lower cost.
Figure 2: Bidirectional Workflow Construction and In-Flow Optimization. Our bidirectional decomposition-aware workflow construction (§ 3.1 ) builds a workflow W at the proper granularity with minimal cost Cmin in two directions: top-down decomposition and bottom-up coalescing (Alg. 3 ). In-flow dynamic optimization (§ 3.2 , Alg. 4 ) then optimizes W by retrospectively applying the cost matrix to construct the credit matrix, fixing faulty intermediate steps in flow without reference answer. The constructed workflow is executed in the sandbox , where InFlowOp optimizes it as the execution proceeds (Alg. 1 ).
Method
Doc
Fin
Chart
Math
Phys
All
Δacc
Baseline: Single LLM ( no tools or skills )
Qwen3.5-4B
5.64
0.52
2.15
0.00
0.00
1.74
–
Qwen3.5-9B
5.95
1.64
4.58
3.31
0.54
3.23
–
GPT-5-mini
4.02
3.34
15.75
10.05
11.79
8.85
–
GPT-5.4-mini
9.67
1.97
17.12
12.32
10.82
10.59
–
GPT-5.4
10.43
2.54
14.07
42.73
13.02
18.62
–
Table 1: Main Results. Mean accuracy (%) per domain on Braid , averaged over the arms each domain holds and over both complexity levels (§ 4 ). All (%) averages over the 16 arms of these five domains.
Figure 3: Evaluation on All Eight Domains of Braid . Mean accuracy (%) per domain for the three models evaluated on the whole benchmark. Each row shows a domain under the two baselines and InFlowOp (§ 5.1 ). The horizontal axis is accuracy (%). Tab. 1 covers the five domains every model is evaluated on.
Method
Doc
Fin
Chart
Math
Phys
All
Δacc
Single Agent
11.26
20.27
25.88
30.38
4.81
17.38
–
greedy-search
8.85
13.46
17.63
33.00
4.55
15.49
-1.89
Coalesce
12.19
21.25
38.55
39.15
7.61
22.21
+ 4.83
Coalesce + in-flow
13.47
22.75
39.99
40.86
9.48
23.80
+ 6.42
InFlowOp
14.81
25.20
41.28
42.65
9.80
25.13
+ 7.75
Table 2: What Each Module Contributes. Mean accuracy (%) of Qwen3.5-9B on Braid . InFlowOp is assembled one module at a time in (2)-(4): (1) a workflow built greedily, (2) a workflow built with Coalesce , (3) the same Coalesce -built workflow with in-flow tempering, and (4) the full method (§ 3 , § D.1 , Alg. 1 ).
Figure 4: How Fit Is Scored, and Where a Fault Is Sought. Mean accuracy (%) on Qwen3.5-4B and Qwen3.5-9B . Left : using rubric estimator against an LLM for cost estimation. Right : in-flow optimization looking where a contract breaches, against the same with the global pass enabled beside it (§ D.5 ).
Figure 5: Construction Earns Its Two Freedoms. Mean accuracy (%) over Qwen3.5 ( 4B and 9B ) among static cost matrix, dynamic cost matrix, and static pool.
Figure 6: A Profile Is Worth What It Remembers. Mean accuracy (%) over an agent scored (1) on its card alone, (2) on a profile evolving within a run, and (3) on a profile evolving across runs, keeping either successful ones only or all of them.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: A Workflow Earns Its Necessity Only Where One Agent Falls Short. The shaded band marks the gap between the two solvers in accuracy (left) and in cost (right). Cost is counted in units of one single-agent run on the original data, and on Braid the two solvers run under matched budgets. Both use gpt-5-mini over 150 samples of SlideVQA ( doc variant of Braid ), Distributed at complexity level k=5 .
Original Dataset
Braid
Composition
Task Type
Data Type
k=3
k=5
Document
MMLongBench-Doc ( Ma et al., 2024 )
MMLongBench-Doc
Distributed
Mixed
Document
426
175
MMLongBench-Doc ( Ma et al., 2024 )
MMLongBench-Doc
Anchored
Mixed
Document
498
305
LongDocURL ( Deng et al., 2025 )
LongDocURL
Distributed
Mixed
Document
441
340
LongDocURL ( Deng et al., 2025 )
LongDocURL
Anchored
Mixed
Document
399
325
Slide
Appendix
Table 3: Data Statistics. Braid as built: 19 evaluation-only arms over 8 domains, with questions each yields at two complexity levels k=3 and k=5 . Original Dataset shows where a Braid task is adapted from. For SlideVQA, we construct two variants: (1) SlideVQA-Doc : providing slides as a document, and (2) SlideVQA-Img : providing slides as individual images. Task Type is the answer type a Braid sample’s k questions demand, and Data Type is the type of sources provided beside input questions. A level is absent where the dataset yields too few questions ( nq<80 ) to evaluate on, and Anchored is absent where no source of that dataset carries k questions to anchor.
Figure 8: The Worth of A Step Is Settled Downstream. (a) Before tempering: the output meets its declared condition and gives two images. However, neither carries the information the subtask demands, and the flow reaches the wrong final answer. (b) After tempering: the output that violates the condition confirms that no image qualifies, and the flow reaches the correct final answer. However, if scored against its own condition, the step that helped is recorded as the step that failed.
Figure 9: What Each Method Spends, and What It Earns. Mean accuracy (%) over the 16 arms of Tab. 1 ( left ), and the tokens a solver spends per task over the same arms ( right , logarithmic), with one color per backbone. S-LLM and S-Agent are the single-LLM and single-agent baselines (§ 5.1 ). Best-of-N repeats workflow construction and execution N=4 times and keeps its best outcome. Both costs are stated as a multiple of one single-agent run on the same backbone. Each method’s turn cost is given beneath its name.
Parameter
Definition
Configuration
Decomposition
Control-flow primitives
The structures a workflow may use beyond plain sequencing (§ D.2 )
And , Repeat , Select , Or
Reliability–latency weight β
What a unit of latency is worth against a unit of reliability in the workflow cost (Eq. 21 )
0.25
Creation penalty γ
What creating one agent costs the same objective (Eq. 21 )
1.0
Competence bar Cˉ
The cost an agent clears to count as covering an atom (Eq. 1 )
0.7
Agent creation
Whether a new agent may be created where no pooled agent covers a subtask
on
Appendix
Table 4: Experiment Configurations. We summarize our experiment configurations for core parameters (§ 5 ), and an ablation arm is studied against the configuration stated here.
Large Language Model (LLM)-based multi-agent systems are increasingly powerful, but current agentic workflow optimization paradigms make an unsatisfying trade-off. Task-level methods spend substantial offline compute yet deploy only a single workflow, leaving complementary candidates unused, while query-level methods synthesize a new workflow per query at substantial inference cost. Our motivating analysis shows these paradigms are more complementary than competing: workflows discovered during offline search often solve different subsets of queries, and many queries handled by expensive query-level generation can already be solved by cheaper precomputed workflows. This suggests a different objective: rather than searching for one universally best workflow or regenerating one per instance, we should build a compact bank of reusable, complementary workflows and select among them adaptively at inference time. Doing so requires solving three coupled problems: generating complementary rather than redundant candidates, compressing them into a small deployable portfolio, and assigning each query to the right workflow under a performance-cost trade-off. To this end, we present FlowBank, a three-stage framework for portfolio-based agentic workflow optimization. Diversifying proposes DiverseFlow to steer search toward under-covered queries and produce a high-coverage candidate pool. Curating proposes CuraFlow to compress this pool into a compact portfolio with minimal redundancy. Matching casts deployment as edge-value prediction on a query-workflow bipartite graph and routes each incoming query to the portfolio member with the best predicted utility. Across five benchmarks, FlowBank achieves the highest average score among the evaluated methods while remaining cost-competitive, improving over the strongest automated and handcrafted baselines by 4.26% and 14.92% relative, respectively.
Tackling complex real-world tasks can exceed the capabilities of a single large language model (LLM), motivating the use of multi-agent workflows that coordinate specialized agents to work together on these tasks. Recent methods train LLMs to construct better workflows from execution outcomes, but they optimize only the workflow generator, while the other agents that build or execute each workflow remain fixed even though every outcome depends on all of them. However, extending training beyond the generator is challenging: the agents are coupled, and a workflow's outcome is a single sparse score that cannot tell which agent causes a failure. We propose FloWright, which leverages the workflow as a harness to optimize workflows. By introducing a hierarchical, structure-aware reward paradigm, FloWright enables one role to self-evolve and two or more roles to co-evolve, with no additional models, labels, or executions. Considering the limitation that workflows are commonly trained and evaluated on data that a single agent can already handle, we further propose DataWright, an adaptive data hardening approach that converts existing datasets into workflow-level tasks with increased difficulty. Across document, slide, chart, code, math, and finance tasks, small open models trained with FloWright achieve improved performance by up to +7.41%, with co-evolving (+5.03%) more roles gaining more than optimizing one of them alone (+2.83%). Our project page: https://xhguo7.github.io/FloWright/.
Xuehang Guo, Haoyu Wang, Haifeng Chen +3
William & Mary · NEC Corporation of America · University of Illinois Urbana-Champaign
Multi-agent systems provide a powerful way to extend large language models (LLMs) by decomposing a complex task into specialized subtasks handled by different agents. However, their performance is often hindered by error propagation, arising from suboptimal workflow design or inaccurate agent outputs, which can propagate through the agent collaboration process and degrade final results. To address the challenges, we present MANGO (Multi-Agent Network Gradient Optimization), a data-driven framework that organizes and refines agent collaboration via a flow network constructed from past successful workflows. MANGO integrates reinforcement learning and textual gradients to jointly optimize workflow paths and agent behaviors, while a skipping mechanism prevents redundant updates to well-optimized agents for improving efficiency. Extensive experiments on seven benchmarks show that MANGO achieves up to 12.8% performance improvement over state-of-the-art baselines, enhances efficiency by 47.4%, and generalizes effectively to unseen domains. Our code and datasets are publicly available at https://github.com/openJiuwen-ai/agent-store/tree/main/community/mango.