Skip to content
AI.info

Research

ANTMAN: Adaptive Need Tracking for Multi-Agent Navigation in Large Information Spaces

Overview Research area: Multi-agent LLM systems for information seeking, long-context reasoning, and structured navigation over large information spaces. Technical level: Advanced. Scope: The paper pr

ANTMAN: Adaptive Need Tracking for Multi-Agent Navigation in Large Information Spaces
arXiv
2609.33326
Published
2026-09-27
Authors
Jerry Wang, Haibo Jin, Xiaopeng Yuan, Peng Kuang, Haohan Wang

AI summary

Overview

Research area: Multi-agent LLM systems for information seeking, long-context reasoning, and structured navigation over large information spaces.

Technical level: Advanced.

Scope: The paper proposes ANTMAN, a coordination framework that replaces static input-partitioning with a revisable "Need Graph" of unresolved information needs, and evaluates it on multi-document QA, controlled long-context scaling, and repository/web navigation benchmarks.

What This Paper Is About

Information-seeking agents increasingly face information spaces too large to process exhaustively. Existing multi-agent long-context systems, such as LongAgent, split the input into fixed-size chunks and assign one worker per chunk, which means the number of active workers grows with how the space is partitioned rather than with what the query actually needs. ANTMAN's goal is to make the evolving set of unresolved information needs itself the runtime control state, so that coordination grows with query demand rather than with information-space size.

Key Contributions

  1. A revisable Need Graph as runtime control state. ANTMAN maintains a graph of information needs that tracks resolution status, accumulated evidence, prior attempt history, and progress. This state governs worker activation, routing, and task-local recovery. Replacing this explicit state with graph-free adaptive replanning reduces performance by 23.82% at 512K and 16.67% on GAIA.

  2. Substrate transfer without substrate-specific redesign. The same need-conditioned mechanism operates across document collections, long-context corpora, software repositories, and web-and-tool environments. ANTMAN outperforms the strongest non-benchmark-specific baselines by 15.3% on RepoProbe, 18.4% on SWE-QA-Pro, and 27.9% on GAIA.

  3. Effectiveness with smaller worker models. With Qwen3-8B used for all non-orchestrator components (ANTMAN-H), the system retains 89.2%, 93.5%, and 98.3% of full ANTMAN's performance on RepoProbe, SWE-QA-Pro, and GAIA, respectively.

  4. Selective coordination with positional robustness. Across Early, Middle, and Late evidence placements, ANTMAN has a middle-position gap of only +.015, compared with -.181 for direct full-context inference.

Main Findings

  • Coordination decouples from space size: Under a 16× increase in searchable context, ANTMAN increases active coordination by only 1.23×, compared with more than 15× for partition-driven baselines (LongAgent and CoA), with correspondingly slower growth in model calls and inference cost.

  • Coordination follows demand: Holding the searchable space fixed at 512K and raising required evidence from 1 to 4 to 16 units raises active coordination from 5.0 to 8.0 workers and model calls from 35.5 to 86.2 per query. ANTMAN recovered all required evidence and answered all evaluated queries correctly (100% complete evidence, 100% answer accuracy) across all three demand levels.

  • Strong multi-document QA performance: Across HotpotQA, 2WikiMultiHopQA, and MuSiQue (30 questions each, nine methods, 810 predictions), ANTMAN achieves macro averages of .711 EM and .814 F1, the strongest aggregate result. ANTMAN-H reaches .622 EM and .763 F1, remaining competitive despite smaller execution models.

  • Structured navigation results: ANTMAN scores 38.21 on RepoProbe-Python, 79.78 on SWE-QA-Pro, and 58.3 overall on GAIA-Text-103. ANTMAN-H scores 34.10, 74.60, and 57.3 respectively. Both exceed general tool-using agents (ReAct: 25.86, 64.92, 34.0; OWL on GAIA: 40.8) and repository-specific agents (ReAct + RepoGraph: 25.71, 67.38; RepoDistill: 28.27, 65.10). The benchmark-specific SWE-QA-Pro Agent scores 77.58.

  • Difficulty-dependent gap on GAIA: Most of the ANTMAN-H gap on GAIA comes from Level 3 (25.0 for ANTMAN-H vs 41.7 for ANTMAN), while ANTMAN-H achieves 55.8 on Level 2 versus ANTMAN's 53.8.

  • Each runtime mechanism contributes differently: Recovery matters most in controlled scaling (removing it drops 512K F1 from 84.03 to 60.12, a 28.45% relative drop), adaptive rerouting matters most in repository navigation (removing it drops SWE-QA-Pro from 81.43 to 72.13, an 11.42% relative drop), and rerouting has no effect on GAIA (60.00, 0.00% change) because GAIA uses a single territory-backed worker.

  • Need revision matters most on GAIA: Removing need revision drops GAIA accuracy from 60.00 to 42.50, a 29.17% relative drop.

  • Positional robustness: ANTMAN's middle-position gap is +.015 versus -.181 for Full Context, which shows the classical "lost in the middle" pattern that ANTMAN avoids.

  • Theoretical bound: Active coordination satisfies |A(q)| ≤ min{M, ρH(q)}, where M is the number of available territories, H(q) is the number of realized information needs, and ρ is the maximum number of routing attempts per need.

Methodology in Plain English

ANTMAN is organized around three pieces.

First, a static substrate map partitions the information space into bounded searchable regions called territories, each described by a lightweight "WorkerCard" summarizing its scope. This map is fixed and defines which workers are available, not which ones must participate.

Second, a revisable Need Graph is initialized from the query and updated after every worker report. Each need node records its resolution status, accumulated evidence, prior attempts, and progress. Updates can resolve a need, keep it unresolved, reframe it, or introduce new needs and dependencies revealed by new evidence.

Third, need-conditioned coordination selects an unresolved need, routes it to a territory-backed worker using the current graph and WorkerCards, and executes bounded local search. If repeated attempts stall, the system applies task-local recovery by reframing the need, rerouting it to another territory, or invoking a fallback path. Routing uses a two-stage procedure: a non-LLM retrieval stage ranks WorkerCards with sparse lexical matching, exact-term matching, dense similarity, and reciprocal-rank fusion, after which the orchestrator decides whether and where to dispatch. Once needs are resolved, the attached evidence is used for final synthesis.

Evaluation spans three settings: standard multi-document QA (HotpotQA, 2WikiMultiHopQA, MuSiQue; 30 questions per benchmark across nine methods), controlled scaling using the multi-needle Needle-in-a-Haystack PLUS setting from LongAgent (10 questions at Early, Middle, and Late positions over 32K, 64K, 128K, and 512K contexts), and realistic navigation (RepoProbe-Python with 108 questions across eight large Python repositories, a frozen 80-question subset of SWE-QA-Pro, and the fixed text-only GAIA-Text-103 subset). Baselines and ANTMAN use GPT-4.1 throughout; ANTMAN-H keeps the GPT-4.1 orchestrator but uses Qwen3-8B for all non-orchestrator components.

Why This Matters

The paper argues that coordination need not mirror the structure of the information space. By making unresolved needs the basis for runtime control, ANTMAN separates available search capacity from the computation actually activated for a query — a design distinction the authors position against partition-driven systems such as LongAgent and CoA, and against adaptive retrieval and adaptive multi-agent systems that change either what information is accessed or how computation is allocated, but not both through a shared evolving need state.

Real-world applications:

  • Repository-level codebase question answering, where an agent must locate and integrate evidence across many files in large repositories without reading everything.
  • Long-document enterprise or legal analysis, where evidence is sparse relative to the corpus and its position in the context affects reliability.
  • Tool-using research assistants operating over heterogeneous web and tool environments, as tested on GAIA-Text-103.
  • Cost-sensitive deployment, where inference cost and model-call volume must stay bounded as the amount of indexed information grows, and where smaller worker models may be substituted for larger ones.

Industry relevance: The results on latency and cost proxies (model calls and dollar cost per query), plus the demonstration that an 8B worker model can retain most of the performance of larger workers, speak directly to production-scale deployments where scaling the corpus should not proportionally scale inference spend.

Future Directions

  • Extending ANTMAN to richer information substrates beyond the document, repository, and web-and-tool settings tested.
  • Developing more transferable representations of information needs that do not require substrate-specific design.
  • Refining the policies for revising and routing needs as search progresses.
  • Investigating why the ANTMAN-H degradation concentrates on harder tasks, particularly GAIA Level 3, and whether smaller worker models can be made more effective on those cases.

Target Audience

Researchers and engineers working on multi-agent LLM systems, long-context reasoning, retrieval-augmented generation, and agentic code or web navigation. It is most useful to readers already familiar with retrieval baselines (DPR, BM25), iterative retrieval methods (ReAct, ChainRAG, S2G-RAG), and multi-agent long-context systems (LongAgent, CoA, OWL), and to practitioners deciding how to allocate coordination and inference budget as their information spaces grow.

Authors’ abstract

Information-seeking agents increasingly operate over information spaces that are too large to process exhaustively. Yet many multi-agent systems organize computation around static partitions of the available space, causing coordination to grow with how information is segmented rather than with what the query still requires. We introduce ANTMAN, an adaptive coordination framework that treats evolving unresolved information needs as the unit of runtime coordination. ANTMAN maintains a revisable Need Graph that tracks unresolved requirements, accumulated evidence, prior attempts, and search progress, and uses this state to control worker selection, routing, and task-local recovery as new evidence is discovered. By separating the coordination policy from substrate-specific search interfaces, the same need-conditioned mechanism can operate across different information spaces. Experiments across multi-document question answering, controlled long-context scaling, and realistic structured navigation show that ANTMAN remains effective across settings, including when execution is delegated to substantially smaller worker models. Under a 16x increase in searchable context, ANTMAN increases active coordination by only 1.23x, compared with more than 15x for partition-driven baselines, while preserving strong answer quality.

Read the original paper