Skip to content
AI.info

Research

SquidAgent: Parallelize Wisely, Coordinate Efficiently

Overview Research area: LLM-based multi-agent systems, specifically execution scheduling and coordination efficiency for agents that decompose tasks into dependency graphs. Technical level: Intermedia

SquidAgent: Parallelize Wisely, Coordinate Efficiently
arXiv
2610.08647
Published
2026-10-06
Authors
Yexiong Lin, Shanshan Ye, Yu Yao, Zhen Fang, Bo Han, Tongliang Liu

AI summary

Overview

  • Research area: LLM-based multi-agent systems, specifically execution scheduling and coordination efficiency for agents that decompose tasks into dependency graphs.
  • Technical level: Intermediate. The framing is intuitive (when is parallelism worth it?), but the paper formalizes the answer with a cost criterion, a safety margin, and a robustness argument in an appendix proof.
  • Scope: The paper identifies two hidden costs of parallel multi-agent execution, proposes a token-based cost criterion for deciding per dependency layer whether to run serially or in parallel, and evaluates the resulting framework (SquidAgent) on 9 tasks against 7 baselines.

What This Paper Is About

LLM agents can solve complex multi-step tasks, and in principle splitting work across parallel agents should give large speedups. In practice, existing parallel multi-agent systems often run slower than a single serial agent, because parallel workers pay hidden costs: they redundantly rebuild context the orchestrator already has ("re-exploration") and their independently produced outputs must be reconciled later ("alignment"). The paper's goal is to make this trade-off explicit and computable before execution, so a system parallelizes a layer of work only when the savings from concurrency exceed those hidden costs.

Key Contributions

  1. Hidden-cost analysis of parallel multi-agent execution. The authors identify two operationally distinct overheads: re-exploration cost (parallel workers redundantly reconstructing the orchestrator's planning context that a serial agent retains implicitly) and alignment cost (reconciling independently generated outputs into a globally consistent result). They show parallel execution can be counterproductive when these costs outweigh concurrency's benefit.

  2. A token-cost criterion for adaptive parallelization. Since wall-clock time is the natural target but is poorly estimated by LLMs, they use predicted output tokens as a backend-independent proxy. They formalize parallelization as cost-sensitive scheduling, comparing predicted serial cost (sum over subtasks) against predicted parallel cost (critical path plus re-exploration and alignment), computed before execution and applied per topological layer of the task DAG.

  3. The SquidAgent framework. An implementation that estimates token budgets in a single planning pass, initializes workers from a shared orchestrator state to reduce redundant re-exploration, and writes a shared convention block before parallel layers to convert alignment into a bounded upfront cost, with a deterministic scheduler applying the criterion layer by layer.

  4. Empirical validation. Across all 9 evaluation tasks, SquidAgent reports a 2.6× average wall-time speedup and a 2.2× average throughput improvement over Claude Code, and a 2.0× average throughput improvement over the strongest multi-agent baseline, while obtaining the best overall task completion quality.

Main Findings

  • Parallel systems can be slower than serial ones: the paper attributes this to two hidden costs — re-exploration and alignment — that a serial agent avoids because it inherits its own context for free.

  • Always-parallelizing is suboptimal: the criterion makes explicit that savings from replacing a sum with a maximum must exceed the added re-exploration and alignment costs; always-serial execution conversely misses real speedups on layers where parallel cost is lower.

  • Token budgets are predicted far more reliably than wall-clock time: across 59 subtasks from the nine tasks, with each quantity predicted five times and the DAG, model, backend, and prompt structure held fixed, output-token predictions had Spearman ρ = 0.77 and Pearson log-transformed r_log = 0.77 with observations, versus ρ = 0.16 and r_log = 0.11 for wall-clock time.

  • Throughput improvements are large and consistent: SquidAgent's overall deliverable throughput was 38.1 ± 12.8 words/s versus 17.2 ± 8.3 for Claude Code and 19.2 ± 5.5 for AgentConductor (the strongest multi-agent baseline). Other baselines: MacNet 15.9 ± 5.1, SeqCV 15.3 ± 7.0, MetaGPT 13.6 ± 4.9, Flow 13.4 ± 6.6, AFlow 11.6 ± 5.0.

  • Gains are largest on big tasks: SquidAgent improved throughput over the strongest baseline by 2.8× on ArcadeBox and 2.3× on SlideKit; on medium tasks the absolute gains were smaller because there are fewer independent subtasks, though SquidAgent still had the best throughput on every task.

  • Quality was preserved or improved: SquidAgent scored 98.2 ± 2.1% overall on task-specific binary rubrics (21–37 items per task, scored YES/NO by Claude Opus 4.6), reaching 100% on 5 of 9 tasks. Claude Code scored 97.1 ± 3.1%, SeqCV 93.3 ± 5.5%, MetaGPT 84.4 ± 28.0%, AgentConductor 81.6 ± 23.3%, Flow 81.2 ± 27.2%, AFlow 68.7 ± 26.2%, and MacNet 62.1 ± 22.1%.

  • All three components matter, and are not additive: on four tasks, mean throughput was 42.93 words/s with the full method, 29.14 without scheduling (a 32.1% decrease), 30.88 without session forking, and 32.63 without convention planning.

  • The safety margin α controls whether borderline layers run in parallel: on MathRef (a 2-layer DAG with 6 chapters in Layer 0 and 4 cross-referencing appendices in Layer 1), at α ≤ 1.8 the scheduler parallelized Layer 1 yielding 34–41 words/s, while at α ≥ 2.0 it serialized Layer 1, raising throughput to 45–51 words/s. The transition occurs between α = 1.8 and α = 2.0, and the default is α = 1.9.

  • Ratio estimates track reality: estimated cost ratios clustered near the diagonal against actual post-execution ratios across 14 multi-task dependency layers from the 9 tasks.

Methodology in Plain English

The orchestrator (planning agent) takes a user request and, in one response, produces three things: a DAG of subtasks, a predicted output-token budget for each subtask, and a predicted alignment-token budget for each dependency layer. No extra LLM calls are needed for these estimates.

The DAG is split into topological layers, where all dependencies of a task are satisfied by earlier layers, so tasks within a layer are eligible to run concurrently. For each layer, the scheduler compares the sum of its subtask token budgets (serial cost) against the slowest subtask's budget plus the layer's alignment budget (parallel cost). Because these are noisy estimates, parallel is chosen only when the serial-to-parallel ratio exceeds a safety margin α; layers with a ratio in (1, α] run serially. The authors show that with bounded ratio-estimation error ε, an ε < α − 1 prevents incorrectly parallelizing any layer whose true ratio is ≤ 1, at the cost of forgoing modest parallel gains.

Two design choices shrink the hidden costs themselves. Workers "fork" the orchestrator's session, inheriting prior reasoning, decisions, constraints, and partial plans, which lets the authors treat re-exploration cost as approximately zero. Before a parallel layer, the orchestrator writes a layer-specific convention block fixing notation, naming, interfaces, formatting, and cross-references, turning alignment from post-hoc repair (hard to predict) into an upfront planning cost that can be estimated.

Experiments use Claude Sonnet (claude-sonnet-4-6) with thinking effort set to medium for all methods, identical task specifications, and the same tool-access setting, with α = 1.9. The 9 tasks span code generation, technical writing, and structured planning. Heavy tasks contain 8–27 files with 300–900 lines per file (PixelCraft, ShopFlow, CompressKit, ArcadeBox, SlideKit, ClimateAnalysis); medium tasks contain 6–10 files with 150–600 lines per file (LinAlgBook, MathRef, CloudDocs). Throughput is words/s of delivered output excluding intermediate artifacts, wall time runs from first model invocation to final deliverable, and quality is judged by task-specific binary rubrics.

Why This Matters

Impact on research: The paper reframes multi-agent parallelism as a cost-sensitive scheduling decision rather than a default, giving a concrete decision rule and naming two overheads (re-exploration, alignment) that prior fixed-policy or heuristic systems leave implicit. It also contributes a measurement result — that LLMs estimate their own output length far better than their own execution time — which is relevant beyond this system.

Real-world applications (as grounded in the paper's task suite):

  • Software development: full-stack applications with backend APIs, authentication, admin logic, and frontend pages that must stay consistent (ShopFlow), and modular toolkits with shared data-format constraints (CompressKit).
  • Game and media module development: loosely coupled game systems in Python/Pygame (PixelCraft) and collections of mostly independent HTML5 Canvas games (ArcadeBox).
  • Long-form document authoring: ML workshop slide decks and handouts requiring cross-document consistency (SlideKit), and analysis scripts plus reports requiring computed results to match written summaries (ClimateAnalysis).
  • Structured technical reference material: linear algebra tutorials, tightly cross-referenced discrete mathematics references, and API documentation with cross-cutting guides (LinAlgBook, MathRef, CloudDocs).

Industry relevance: Any product that runs fleets of LLM agents on multi-step work — coding assistants, document generation pipelines, workflow automation — pays for wasted parallel workers and reconciliation. A decision rule that avoids parallelizing when it does not pay off, and that reports 2.2× mean throughput over Claude Code and 2.0× over the strongest multi-agent baseline at higher judged quality, targets directly the latency and cost metrics such products are measured on. The paper releases code to support reuse.

Future Directions

  • Adaptive safety margins: the fixed margin α = 1.9 does not adapt to uncertainty in token estimates; a margin that responds to estimate confidence is an open direction.
  • Costs beyond token generation: the token proxy does not capture latency from tool execution, retrieval, or external API calls, so extending the criterion to latency-inclusive costs remains open.
  • Finer-grained scheduling: SquidAgent makes a single serial-or-parallel decision per layer, even when some tasks in that layer would benefit from parallelism while others would be better run serially.
  • Beyond ratio estimation error: the analysis bounds robustness under an assumed error bound ε on the ratio estimate; how estimate errors behave across more models, backends, and task types (the paper also reports transfer-to-external-task experiments in an appendix) is a natural extension.

Target Audience

Researchers and engineers working on LLM agent systems, multi-agent orchestration, and inference-time efficiency; practitioners building coding assistants, document generation pipelines, or workflow automation on top of agent frameworks; and readers interested in scheduling and cost-benefit decision rules for delegated AI work. Readers wanting only the conceptual lesson — that parallelism has re-exploration and alignment costs and should be chosen layer by layer — can follow the main text without the appendix formalization.

Authors’ abstract

LLM-based agents solve complex multi-step tasks, but sequential execution incurs substantial latency. In principle, parallelizing work across multiple agents should yield near-linear speedups. Yet existing parallel multi-agent systems often run slower than a single-agent baseline. We attribute this gap to two hidden costs that parallel execution incurs but a serial agent avoids. First, there is a re-exploration cost: redundant effort spent by parallel workers reconstructing context that the orchestrator already possesses, such as prior decisions, that would otherwise be inherited implicitly in a serial execution. Second, there is an alignment cost: the overhead required to reconcile inconsistencies across independently generated outputs. We thus derive a principled decision criterion: a layer should be parallelized only when its critical-path cost, plus re-exploration and alignment overheads, is lower than the corresponding serial cost. While this criterion is naturally expressed in wall-clock time, we observe that LLMs are poorly calibrated when asked to estimate task duration. To address this, we instead measure cost in predicted output tokens, which we empirically find LLMs can estimate substantially more reliably than wall-clock time. Building on this token-based criterion, we propose SquidAgent. It estimates all token budgets in a single planning step, forks each worker directly from the orchestrator's session to eliminate re-exploration cost, and replaces post-hoc reconciliation with a pre-generated shared convention block that converts alignment into a bounded upfront cost. A deterministic scheduler then applies the criterion layer by layer. Empirically, SquidAgent achieves a 2.2$\times$ mean throughput improvement and a 2.6$\times$ mean wall-time speedup over Claude Code, and a 2.0$\times$ throughput improvement over the strongest multi-agent baseline.

Read the original paper