Skip to content
AI.info

Research

DAGent: Evaluate-then-Grow Planning for Deep Research Agents

DAGent: Evaluate-then-Grow Planning for Deep Research Agents Overview Research area: Natural Language Processing / LLM agent architectures, specifically planning and reinforcement learning for deep re

DAGent: Evaluate-then-Grow Planning for Deep Research Agents
arXiv
2609.39154
Published
2026-09-30
Authors
Hanwen Liu, Yuanfu Sun, Qiaoyu Tan

AI summary

DAGent: Evaluate-then-Grow Planning for Deep Research Agents

Overview

Research area: Natural Language Processing / LLM agent architectures, specifically planning and reinforcement learning for deep research agents that decompose tasks into directed acyclic graphs (DAGs).

Technical level: Advanced. The paper assumes familiarity with multi-agent orchestration, ReAct loops, context management strategies, and group-based policy-gradient RL (GRPO, DAPO).

Scope: The paper proposes DAGent, a DAG-based multi-agent framework that replaces upfront planning with incremental, evidence-conditioned graph growth, plus DAGRPO, a DAG-conditioned RL recipe, and evaluates both on three deep research benchmarks across multiple backbones.

What This Paper Is About

Deep research tasks require an AI agent to explore large knowledge spaces, gather evidence from many sources, and adapt its plan as intermediate findings appear. Existing DAG-based agents plan the whole task graph before executing anything and then repair it after failures — a pattern the authors call Plan-then-Patch — which commits most heavily at the moment the system knows least. DAGent instead grows the task graph one batch of nodes at a time, conditioning each expansion on structured confidence and uncertainty signals from already-completed nodes.

Key Contributions

  1. Evaluate-then-Grow planning. An incremental DAG planning paradigm in which an Orchestrator grows the task graph batch by batch, driven by structured per-node feedback (sub-task outcome, rationale, and reliability) rather than an upfront commitment followed by reactive edits. The authors state DAGent is the first deep research agent to condition each planning decision on structured per-node feedback.

  2. DAGRPO. A DAG-conditioned reinforcement learning variant that adds two structural signals on top of a GRPO-style recipe: topology-conditioned credit on Executor rollouts (based on whether their outputs feed the final synthesis) and a structural compliance regularization on Orchestrator plans that fail validation.

  3. Broad empirical advantage. The lead is shown across BrowseComp-Plus, GAIA, and xbench-DeepSearch; across four open-source backbones from four vendors; and when scaled to GPT-5 at 327K context. DAGRPO is shown to add gains over an outcome-only GRPO baseline at the same compute budget.

  4. Efficiency finding. A same-architecture comparison shows incremental, evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart.

Main Findings

  • Training-free lead at every Qwen3 scale. At Qwen3-32B, DAGent reaches 47.3 / 55.3 / 65.0 Pass@1 on BrowseComp-Plus / GAIA / xbench-DeepSearch, which is 4.0 / 3.8 / 2.0 points above FlowSearch (the strongest baseline at that scale) and 6.0 / 10.6 / 7.0 points above the strongest of the remaining baselines. The lead persists at 8B (+5.3 / +7.8 / +5.0) and at 235B-A22B (+5.3 / +5.8 / +2.0), matching the numbers quoted in the abstract. On every difficulty split DAGent matches or exceeds every baseline, with two ties at 235B-A22B; relative gains grow with difficulty at 8B and 32B.

  • DAGRPO gain isolates to structural signals. DAGRPO raises Qwen3-8B DAGent from 40.0 / 46.6 / 60.0 in training-free mode to 49.6 / 53.4 / 65.7 over three seeds, exceeding the GRPO-DAGent baseline (46.0 / 50.8 / 63.0) by 3.6 / 2.6 / 2.7 points, or 3.0 average points. The same outcome-only GRPO recipe improves each trainable baseline by 2.9 to 10.7 points across the three benchmarks, so the authors attribute the further gain to the structural signals rather than the training budget.

  • Replication across vendors and frontier scale. DAGent holds the per-backbone top position across Qwen3-32B, Seed-OSS-36B, GLM-4-32B, and Nemotron-3-Nano-30B. At GPT-5 with 327K context it leads the strongest reported baseline by 5.7 / 2.5 / 2.0 points on BrowseComp-Plus / GAIA / xbench-DeepSearch, with GAIA evaluated on its full 165-task validation set using multi-modal tools.

  • Planning paradigm matters more than context components. In the Qwen3-32B ablation, replacing Evaluate-then-Grow with a Plan-then-Patch variant costs 5.3 / 4.8 / 4.0 points; collapsing further to a single upfront plan with no replanning costs another 8.7 / 7.8 / 6.0, totaling 14.0 / 12.6 / 10.0. Among context components the drop ordering is consistent across benchmarks: removing Selective Propagation costs 12.6 / 10.6 / 8.0, removing QueryDoc costs 8.0 / 7.7 / 5.0, and removing InteractionTranscript costs 4.0 / 2.9 / 2.0 (this last row is also the RecallTool ablation).

  • Both DAGRPO signals contribute; the off-chain coefficient peaks at 0.5. Removing topology-conditioned credit costs 2.3 / 1.9 / 1.7 points; removing structural compliance regularization costs 1.6 / 1.0 / 0.7 points. The ablation rows use a single seed while the DAGRPO row is a three-seed mean with std 1.0 / 1.0 / 1.5, so the smaller per-benchmark differences are within seed noise; the authors rest the claim on the consistent sign of every per-benchmark difference and the Overall differences of −2.0 and −1.1. The α sweep peaks at 0.5 on every benchmark; α = 0 pushes performance below the GRPO baseline, while α = 0.25 retains 75 / 62 / 37% of the gain over the GRPO baseline.

  • DAGent uses fewer resources than its Plan-then-Patch counterpart. Median per-task tool-call-to-step ratios are 1.6, 1.1, and 1.2 on BrowseComp-Plus, GAIA, and xbench-DeepSearch. Mean per-task footprints are 42.7 / 31.2 / 26.8 steps, 67.6 / 36.8 / 32.9 tool calls, and 1.20M / 0.66M / 0.44M total input-plus-output tokens. The Plan-then-Patch variant raises the off-chain Executor ratio from 0.25 / 0.20 / 0.18 to 0.40 / 0.30 / 0.25 and adds 40 / 29 / 18% more tokens, 28 / 21 / 21% more execution steps, 35 / 27 / 21% more external tool calls, and 35 / 27 / 23% more wall-clock time per task, while still trailing DAGent by 5.3 / 4.8 / 4.0 Pass@1 points.

  • RecallTool is a fallback, not a default. In the Qwen3-32B runs it is invoked in 22.0 / 13.6 / 11.0% of tasks on BrowseComp-Plus / GAIA / xbench-DeepSearch.

Methodology in Plain English

The problem with planning everything upfront. In a Plan-then-Patch system, the agent writes a full task decomposition before running any sub-task, then edits that graph when something fails or evidence is missing. The authors argue this front-loads the biggest commitment when understanding is weakest, and later revisions waste computation on branches that should never have been planned.

Growing the graph instead. DAGent starts with an empty graph. An Orchestrator examines the most recently executed batch of nodes and assigns each one a status: Success (confident evidence), Uncertain (needs verification), or Not Found (valid execution but no relevant evidence). Nodes marked Uncertain or Not Found spawn "refine" nodes, with each line capped at three attempts (one initial execution plus two refinements). New nodes are added with their descriptions, prompts, and dependency edges; nodes in a batch with no shared dependencies run in parallel. The loop ends when the Orchestrator schedules an answer-type node, which must appear alone in its batch so it can read all the evidence selected for it.

Two context layers. Each node produces a compact QueryDoc containing the answer, a brief explanation, key execution steps, a confidence value in [0, 1], and a list of uncertainties. It also preserves the full InteractionTranscript of its actions and observations. By default only the QueryDocs of direct dependencies propagate downstream — a policy called Selective Propagation, which bounds a node's context by its fan-in rather than graph depth or total node count. When the compact summary is insufficient, the Executor can call RecallTool to pull a focused evidence snippet from a direct dependency's full transcript.

Teaching the model with structure. Because the graph is append-only, a node's ancestry in the final graph matches what synthesis actually read. DAGRPO exploits this: Executor rollouts on the answer-inclusive closure (the answer node plus all its ancestors) keep the full trajectory reward, while off-chain rollouts are attenuated by a coefficient α. Orchestrator turns that violate structural constraints — missing status assignments, unmet refine budgets, an answer node not scheduled alone, duplicate prompts, unparseable JSON with bad dependency references — receive a negative regularization term.

How it was tested. Benchmarks are BrowseComp-Plus (local dense retrieval over a verified corpus with Qwen3-Embedding-8B, using a 680/150 train/evaluation split balanced across easy, medium, and hard), GAIA (the 103-task text-only validation subset with live Google search via Serper and page extraction via Jina; the full 165-task set only for the GPT-5 327K comparison), and the 2505 release of xbench-DeepSearch (100 Chinese-language tasks). Backbones are Qwen3-8B, Qwen3-32B, and Qwen3-235B-A22B for training-free evaluation and Qwen3-8B for RL training, all with thinking disabled and a 32K-token per-sub-task response budget on top of an 8K-token prompt. RL trains on BrowseComp-Plus with the DAPO asymmetric clip, α = 0.5, and the structural compliance regularization; inference uses greedy decoding and Pass@1 scored by a human-calibrated LLM judge.

Why This Matters

The paper's central argument is architectural rather than incremental: how and when an agent commits to a plan may matter as much as how much compute it spends. The ablation supports this, since the Plan-then-Patch-to-single-plan collapse costs far more accuracy than removing any individual context component, and the same-architecture efficiency comparison shows the incremental approach is both more accurate and cheaper per task.

  • Research assistance: literature review, evidence synthesis across dozens or hundreds of sources, and multi-hop fact verification where intermediate findings redirect the investigation.
  • Enterprise and competitive intelligence: long-horizon queries over private corpora plus live web search, where the retrieval backend can be swapped for an internal index.
  • Chinese-language and multilingual research: xbench-DeepSearch is a 100-task Chinese benchmark sharing the Serper/Jina stack with GAIA, indicating the approach is not English-specific.
  • Cost-sensitive agent deployment: the per-task step, tool-call, and token measurements make the case that evidence-conditioned growth avoids redundant branches rather than buying accuracy with more computation.

Industry relevance: the method targets exactly the production concern of deep research agents — runaway token cost and wasted parallel branches. The efficiency results (lower token, tool-call, step, and wall-clock footprints than Plan-then-Patch) and the training recipe (RL on Qwen3-8B, a small open backbone) suggest the gains are reachable without frontier-scale models, though the GPT-5 327K result depends on a closed backbone.

Future Directions

  • Preserving the append-only property. DAGRPO requires that ancestry in the final graph matches what synthesis actually read; the authors note that Plan-then-Patch trajectories do not preserve this because nodes can be deleted or rewired. How to extend structural credit signals to graph-editing agents is left open (the paper's appendix is truncated before this is fully addressed, so no such method is reported).

  • Reducing seed sensitivity in ablations. The structural-signal ablations use a single seed while the DAGRPO row is a three-seed mean, and the per-benchmark deltas for the smaller signal fall within seed noise. Multi-seed verification of both removal experiments would tighten these claims.

  • Scaling the RL recipe. DAGRPO is only trained at the Qwen3-8B scale and only on BrowseComp-Plus, with zero-shot transfer to GAIA and xbench-DeepSearch. Whether the topology-conditioned credit remains effective at 32B or with heavier off-policy data is not reported.

  • Tuning the structural constraints. The regularization involves λ_proc, K_proc, and a three-attempt refine budget. The paper reports the α sweep but does not report a sweep over these other structural hyperparameters.

Target Audience

Researchers and engineers working on LLM agent architectures, multi-agent orchestration, and reinforcement learning for long-horizon tasks will get the most from this paper. It is particularly relevant to practitioners building production deep research or retrieval agents who care about per-task cost, and to RL researchers interested in how recorded trajectory structure can be turned into credit-assignment signals that outcome-only group-based recipes cannot express. Readers without background in GRPO-style policy optimization will find the method section demanding, though the architectural argument in the introduction and ablation study is accessible on its own.

Authors’ abstract

Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting because they support parallel execution and isolate each sub-task within a focused dependency context. Yet existing DAG-based agents instantiate a task-level plan before execution and repair the graph only after failures or missing evidence are observed. This Plan-then-Patch strategy is brittle for deep research: the system commits most strongly when its evidence is weakest, and later revisions waste computation on branches that should not have been planned. We propose DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes. A hierarchical context layer propagates compact QueryDocs by default while preserving full execution traces for on-demand recall. The recorded DAG topology admits structural RL signals that outcome-only recipes cannot define; DAGRPO, a GRPO adaptation, injects topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans. Across BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, and the lead replicates across four open-source backbones and extends to GPT-5 at 327K context. At the Qwen3-8B scale, DAGRPO improves over a same-budget outcome-only GRPO baseline by 3.0 average Pass@1 points. A same-architecture comparison shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart. Code: https://github.com/hanwenliu6825/DAGent

Read the original paper