cs.AIOct 6, 2026

SquidAgent: Parallelize Wisely, Coordinate Efficiently

Authors: Yexiong Lin, Shanshan Ye, Yu Yao, Zhen Fang, Bo Han, Tongliang Liu

Organizations: The University of Sydney · Mohamed bin Zayed University of Artificial Intelligence · University of Technology Sydney · Hong Kong Baptist University

Abstract

LLM-based agents solve complex multi-step tasks, but sequential execution incurs substantial latency. In principle, parallelizing work across multiple agents should yield near-linear speedups. Yet existing parallel multi-agent systems often run slower than a single-agent baseline. We attribute this gap to two hidden costs that parallel execution incurs but a serial agent avoids. First, there is a re-exploration cost: redundant effort spent by parallel workers reconstructing context that the orchestrator already possesses, such as prior decisions, that would otherwise be inherited implicitly in a serial execution. Second, there is an alignment cost: the overhead required to reconcile inconsistencies across independently generated outputs. We thus derive a principled decision criterion: a layer should be parallelized only when its critical-path cost, plus re-exploration and alignment overheads, is lower than the corresponding serial cost. While this criterion is naturally expressed in wall-clock time, we observe that LLMs are poorly calibrated when asked to estimate task duration. To address this, we instead measure cost in predicted output tokens, which we empirically find LLMs can estimate substantially more reliably than wall-clock time. Building on this token-based criterion, we propose SquidAgent. It estimates all token budgets in a single planning step, forks each worker directly from the orchestrator's session to eliminate re-exploration cost, and replaces post-hoc reconciliation with a pre-generated shared convention block that converts alignment into a bounded upfront cost. A deterministic scheduler then applies the criterion layer by layer. Empirically, SquidAgent achieves a 2.2×\times mean throughput improvement and a 2.6×\times mean wall-time speedup over Claude Code, and a 2.0×\times throughput improvement over the strongest multi-agent baseline.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems

    Aug 6, 2026Zihan Xu, Haolin Tian, Hai JiangMulti-Agent Large Language Model SystemsAgentic Inference

  2. When Parallelism Pays Off: Cohesion-Aware Task Partitioning for Multi-Agent Coding

    May 31, 2026Xu Yang, Lunyiu Nie, Ethan Chandra +3Multi-Agent OrchestrationCoding Agents

  3. Response-Conditioned Parallel-to-Sequential Orchestration for Multi-Agent Systems

    May 15, 2026Nurbek Tastan, Alex Iacob, Lorenzo Sani +4Large Language Model Agents