Large language models (LLMs) increasingly construct multi-agent workflows that decompose a complex task and assign specialist agents from a pool. However, building such a workflow well remains challenging: how finely to divide the task, which agent to trust with each subtask, and when to create a new specialist are all critical decisions a workflow constructor needs to settle up front. Thus, whether each subtask succeeds remains unknown until the workflow runs. Yet, improving a workflow is costly. Locating a fault usually requires a reference answer, a graded outcome, or a trained assessor, and the fix is applied to the whole workflow through re-execution, re-search, or retraining. We propose InFlowOp, which prices every decision in one label-free cost that weighs how well an agent's competence meets what a subtask demands against how much that agent takes to run. Before execution, InFlowOp bidirectionally determines the granularity of task decomposition and agent assignment following from the cost rather than from a fixed template. During execution, InFlowOp corrects a fault with the cheapest move via the same cost that serves the workflow both as it is built and as it runs. Facing the workflow-level evaluation challenge, we introduce Braid, a benchmark whose tasks require multi-agent coordination beyond single-agent capability. Across various domains and backbones, InFlowOp outperforms single agent baselines by up to +11.97%, achieving +9.64% with in-flow optimization. Our project page: https://xhguo7.github.io/InFlowOp/.
Figures & tables
Figure 1: InFlowOp for In-Flow Optimization. Complex tasks challenge single agents with long-context reasoning, while multi-agent systems introduce execution failures and coordination overhead. InFlowOp optimize agent behaviors in-flow, enabling efficient collaboration at lower cost.
Figure 2: Bidirectional Workflow Construction and In-Flow Optimization. Our bidirectional decomposition-aware workflow construction (§ 3.1 ) builds a workflow W at the proper granularity with minimal cost Cmin in two directions: top-down decomposition and bottom-up coalescing (Alg. 3 ). In-flow dynamic optimization (§ 3.2 , Alg. 4 ) then optimizes W by retrospectively applying the cost matrix to construct the credit matrix, fixing faulty intermediate steps in flow without reference answer. The constructed workflow is executed in the sandbox , where InFlowOp optimizes it as the execution proceeds (Alg. 1 ).
Method
Doc
Fin
Chart
Math
Phys
All
Δacc
Baseline: Single LLM ( no tools or skills )
Qwen3.5-4B
5.64
0.52
2.15
0.00
0.00
1.74
–
Qwen3.5-9B
5.95
1.64
4.58
3.31
0.54
3.23
–
GPT-5-mini
4.02
3.34
15.75
10.05
11.79
8.85
–
GPT-5.4-mini
9.67
1.97
17.12
12.32
10.82
10.59
–
GPT-5.4
10.43
2.54
14.07
42.73
13.02
18.62
–
Table 1: Main Results. Mean accuracy (%) per domain on Braid , averaged over the arms each domain holds and over both complexity levels (§ 4 ). All (%) averages over the 16 arms of these five domains.
Figure 3: Evaluation on All Eight Domains of Braid . Mean accuracy (%) per domain for the three models evaluated on the whole benchmark. Each row shows a domain under the two baselines and InFlowOp (§ 5.1 ). The horizontal axis is accuracy (%). Tab. 1 covers the five domains every model is evaluated on.
Method
Doc
Fin
Chart
Math
Phys
All
Δacc
Single Agent
11.26
20.27
25.88
30.38
4.81
17.38
–
greedy-search
8.85
13.46
17.63
33.00
4.55
15.49
-1.89
Coalesce
12.19
21.25
38.55
39.15
7.61
22.21
+ 4.83
Coalesce + in-flow
13.47
22.75
39.99
40.86
9.48
23.80
+ 6.42
InFlowOp
14.81
25.20
41.28
42.65
9.80
25.13
+ 7.75
Table 2: What Each Module Contributes. Mean accuracy (%) of Qwen3.5-9B on Braid . InFlowOp is assembled one module at a time in (2)-(4): (1) a workflow built greedily, (2) a workflow built with Coalesce , (3) the same Coalesce -built workflow with in-flow tempering, and (4) the full method (§ 3 , § D.1 , Alg. 1 ).
Figure 4: How Fit Is Scored, and Where a Fault Is Sought. Mean accuracy (%) on Qwen3.5-4B and Qwen3.5-9B . Left : using rubric estimator against an LLM for cost estimation. Right : in-flow optimization looking where a contract breaches, against the same with the global pass enabled beside it (§ D.5 ).
Figure 5: Construction Earns Its Two Freedoms. Mean accuracy (%) over Qwen3.5 ( 4B and 9B ) among static cost matrix, dynamic cost matrix, and static pool.
Figure 6: A Profile Is Worth What It Remembers. Mean accuracy (%) over an agent scored (1) on its card alone, (2) on a profile evolving within a run, and (3) on a profile evolving across runs, keeping either successful ones only or all of them.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: A Workflow Earns Its Necessity Only Where One Agent Falls Short. The shaded band marks the gap between the two solvers in accuracy (left) and in cost (right). Cost is counted in units of one single-agent run on the original data, and on Braid the two solvers run under matched budgets. Both use gpt-5-mini over 150 samples of SlideVQA ( doc variant of Braid ), Distributed at complexity level k=5 .
Original Dataset
Braid
Composition
Task Type
Data Type
k=3
k=5
Document
MMLongBench-Doc ( Ma et al., 2024 )
MMLongBench-Doc
Distributed
Mixed
Document
426
175
MMLongBench-Doc ( Ma et al., 2024 )
MMLongBench-Doc
Anchored
Mixed
Document
498
305
LongDocURL ( Deng et al., 2025 )
LongDocURL
Distributed
Mixed
Document
441
340
LongDocURL ( Deng et al., 2025 )
LongDocURL
Anchored
Mixed
Document
399
325
Slide
Appendix
Table 3: Data Statistics. Braid as built: 19 evaluation-only arms over 8 domains, with questions each yields at two complexity levels k=3 and k=5 . Original Dataset shows where a Braid task is adapted from. For SlideVQA, we construct two variants: (1) SlideVQA-Doc : providing slides as a document, and (2) SlideVQA-Img : providing slides as individual images. Task Type is the answer type a Braid sample’s k questions demand, and Data Type is the type of sources provided beside input questions. A level is absent where the dataset yields too few questions ( nq<80 ) to evaluate on, and Anchored is absent where no source of that dataset carries k questions to anchor.
Figure 8: The Worth of A Step Is Settled Downstream. (a) Before tempering: the output meets its declared condition and gives two images. However, neither carries the information the subtask demands, and the flow reaches the wrong final answer. (b) After tempering: the output that violates the condition confirms that no image qualifies, and the flow reaches the correct final answer. However, if scored against its own condition, the step that helped is recorded as the step that failed.
Figure 9: What Each Method Spends, and What It Earns. Mean accuracy (%) over the 16 arms of Tab. 1 ( left ), and the tokens a solver spends per task over the same arms ( right , logarithmic), with one color per backbone. S-LLM and S-Agent are the single-LLM and single-agent baselines (§ 5.1 ). Best-of-N repeats workflow construction and execution N=4 times and keeps its best outcome. Both costs are stated as a multiple of one single-agent run on the same backbone. Each method’s turn cost is given beneath its name.
Parameter
Definition
Configuration
Decomposition
Control-flow primitives
The structures a workflow may use beyond plain sequencing (§ D.2 )
And , Repeat , Select , Or
Reliability–latency weight β
What a unit of latency is worth against a unit of reliability in the workflow cost (Eq. 21 )
0.25
Creation penalty γ
What creating one agent costs the same objective (Eq. 21 )
1.0
Competence bar Cˉ
The cost an agent clears to count as covering an atom (Eq. 1 )
0.7
Agent creation
Whether a new agent may be created where no pooled agent covers a subtask
on
Appendix
Table 4: Experiment Configurations. We summarize our experiment configurations for core parameters (§ 5 ), and an ablation arm is studied against the configuration stated here.