Authors: Yexiong Lin, Shanshan Ye, Yu Yao, Zhen Fang, Bo Han, Tongliang Liu
Organizations: The University of Sydney · Mohamed bin Zayed University of Artificial Intelligence · University of Technology Sydney · Hong Kong Baptist University
LLM-based agents solve complex multi-step tasks, but sequential execution incurs substantial latency. In principle, parallelizing work across multiple agents should yield near-linear speedups. Yet existing parallel multi-agent systems often run slower than a single-agent baseline. We attribute this gap to two hidden costs that parallel execution incurs but a serial agent avoids. First, there is a re-exploration cost: redundant effort spent by parallel workers reconstructing context that the orchestrator already possesses, such as prior decisions, that would otherwise be inherited implicitly in a serial execution. Second, there is an alignment cost: the overhead required to reconcile inconsistencies across independently generated outputs. We thus derive a principled decision criterion: a layer should be parallelized only when its critical-path cost, plus re-exploration and alignment overheads, is lower than the corresponding serial cost. While this criterion is naturally expressed in wall-clock time, we observe that LLMs are poorly calibrated when asked to estimate task duration. To address this, we instead measure cost in predicted output tokens, which we empirically find LLMs can estimate substantially more reliably than wall-clock time. Building on this token-based criterion, we propose SquidAgent. It estimates all token budgets in a single planning step, forks each worker directly from the orchestrator's session to eliminate re-exploration cost, and replaces post-hoc reconciliation with a pre-generated shared convention block that converts alignment into a bounded upfront cost. A deterministic scheduler then applies the criterion layer by layer. Empirically, SquidAgent achieves a 2.2× mean throughput improvement and a 2.6× mean wall-time speedup over Claude Code, and a 2.0× throughput improvement over the strongest multi-agent baseline.
Figures & tables
Figure 1 : Comparison of two execution strategies on a task with four subtasks. (a) Sequential: a single agent executes all tasks one by one with no overhead. (b) Naive parallel: workers run concurrently but incur re-exploration cost (orange: orchestrator planning plus per-worker rediscovery of shared context) and alignment cost (red: fixing conflicts between independently produced outputs), making the total time exceed the sequential baseline.
Heavy
Medium
Method
Pixel.
Shop.
Comp.
Arcade.
Slide.
Climate.
LinAlg.
Math.
Cloud.
Overall
SquidAgent
21.0
35.1
16.1
52.6
47.7
27.7
46.3
48.1
48.2
38.1 ± 12.8
Claude Code
5.0
16.7
3.7
17.5
19.9
13.5
22.6
26.7
29.6
17.2 ± 8.3
SeqCV
5.3
15.6
4.0
14.0
18.1
15.8
21.8
27.9
15.6
15.3 ± 7.0
MetaGPT
12.3
16.3
6.9
8.7
9.7
11.4
20.2
22.4
14.3
13.6 ± 4.9
AFlow
4.2
8.8
3.7
14.8
12.7
9.5
16.8
18.3
15.4
11.6 ± 5.0
Table 1 : Deliverable throughput (words/s) across 9 evaluation tasks. Computed as deliverable output divided by wall time. Higher is better. Bold indicates the best result per task.
Figure 2 : Scheduler sensitivity analysis. (a) Estimated cost ratio ρk versus actual cost ratio measured after execution, across 14 multi-task dependency layers from 9 evaluation tasks. Layer 0 and Layer 1 correspond to the first and second dependency layers in the orchestrator-generated task DAG. (b) Throughput on MathRef under different safety margins α , with the task DAG and token estimates held fixed. For α≥2.0 , Layer 1 changes from parallel to serial execution, which improves throughput because the alignment cost of cross-referencing appendices exceeds the parallelism benefit.
Method
Math.
Cloud.
Slide.
Climate.
Mean
SquidAgent
48.10
48.20
47.70
27.70
42.93
Without scheduling
35.54
31.70
31.37
17.96
29.14
Without session forking
41.22
33.02
30.26
19.04
30.88
Without convention planning
39.90
35.33
32.78
22.50
32.63
Table 3: Component-wise ablation: deliverable throughput (words/s) on four tasks. Higher is better.
Prediction
ρ
rlog
Output tokens
0.77
0.77
Wall-clock time
0.16
0.11
Table 4: Prediction–observation correlations.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Output type
Main coordination challenge
Scale
PixelCraft
Python/Pygame
Independent game modules with shared runtime logic
Large
ShopFlow
Flask + HTML/JS
Backend–frontend consistency and route integration
Large
CompressKit
Python
Modular utilities with shared APIs and tests
Large
ArcadeBox
HTML5 Canvas/JS
Multiple interactive components under one interface
Large
SlideKit
HTML
Consistent slide structure, styling, and navigation
Large
ClimateAnalysis
Python + Markdown
Multi-stage analysis, reporting, and cross-reference consistency
Large
Appendix
Table 5 : Characteristics of the evaluation tasks. Large tasks require substantial software systems or document collections, while medium tasks require smaller but still multi-part outputs with consistency constraints.
Heavy
Medium
Method
Pixel.
Shop.
Comp.
Arcade.
Slide.
Climate.
LinAlg.
Math.
Cloud.
Overall
SquidAgent
580
404
344
439
1022
464
147
566
383
483 ± 239
Claude Code
1999
1000
1790
1105
2297
1000
529
757
614
1232 ± 638
SeqCV
1115
2231
2780
2620
3127
2214
863
2243
3490
2298 ± 860
MetaGPT
3356
2396
3696
4412
3301
1843
479
969
897
2372 ± 1403
AFlow
1353
534
949
395
168
189
307
401
307
511 ± 392
Appendix
Table 6 : Wall time (seconds) across 9 evaluation tasks. Lower is better.
Heavy
Medium
Method
Pixel.
Shop.
Comp.
Arcade.
Slide.
Climate.
LinAlg.
Math.
Cloud.
Overall
SquidAgent
12,180
14,180
5,538
23,091
48,749
12,853
6,806
27,225
18,461
18,787 ± 13,260
Claude Code
9,995
16,700
6,623
19,338
45,710
13,500
11,955
20,212
18,174
18,023 ± 11,328
SeqCV
5,910
34,804
11,120
36,680
56,599
34,981
18,813
62,580
54,444
35,103 ± 20,269
MetaGPT
41,279
39,055
25,502
38,384
32,020
21,010
9,676
21,706
12,827
26,829 ± 11,570
AFlow
5,681
4,696
3,511
5,844
2,133
1,796
5,157
7,332
4,728
4,542 ± 1,790
Appendix
Table 7 : Deliverable output (words) across 9 evaluation tasks. We count only final deliverable files (source code, content documents), excluding intermediate artifacts such as logs, planning documents, and build outputs.
task
Layer
Tasks
Calg(Lk)
Ratio
Decision
MathRef
L0: chapters
6
2500
5.09
Parallel
L1: appendices
4
6000
1.45
Serial
SlideKit
L0: decks
6
4000
4.80
Parallel
L1: handouts
3
3500
2.00
Parallel
CloudDocs
L0: modules
5
3500
3.31
Parallel
L1: guides
3
3000
1.64
Serial
Appendix
Table 8 : Scheduler decisions on multi-layer DAG evaluation tasks. Calg(Lk) is estimated separately for each layer. The scheduler parallelizes a layer only when the ratio exceeds α=1.9 .
Predicted quantity
Kendall’s τb↑
Inverted pairs ↓
Output-token count
0.62[0.30,0.79]
17%[8%,33%]
Wall-clock execution time
0.12
44%[33%,55%]
Appendix
Table 9: Rank-ordering agreement with the corresponding observed quantities. Inverted-pair fractions are computed among non-tied comparisons; brackets report 95% confidence intervals.
Task
Layer
Mean ρk
Range
Decision
MathRef
Chapters
4.64
[4.53,4.73]
Parallel
MathRef
Appendices
1.71
[1.70,1.74]
Serial
SlideKit
Decks
4.03
[3.75,4.17]
Parallel
SlideKit
Handouts
1.71
[1.53,1.88]
Serial
CloudDocs
Modules
3.16
[3.06,3.31]
Parallel
CloudDocs
Guides
1.63
[1.44,1.75]
Serial
Appendix
Table 10: Repeated-planning sensitivity results. Mean and range are across three planning runs; decisions use α=1.9 .
Metric
Method
Website
Game
Beamer
Mean
Wall time ↓
Claude Code
369
698
92
386.3
SquidAgent
259
439
83
260.3
Throughput ↑
Claude Code
19.8
4.0
10.9
11.6
SquidAgent
28.1
6.2
12.0
15.4
Quality ↑
Claude Code
100.0
100.0
97.0
99.0
SquidAgent
100.0
100.0
100.0
100.0
Appendix
Table 11: Preliminary transfer evaluation on three external tasks from Flow. Wall time is in seconds, throughput in words/s, and quality in percent.