Research
Hierarchical Deep Research with Local-Web RAG: Toward Automated System-Level Materials Discovery
Overview Research area: Large language model (LLM) agents for scientific discovery, specifically hierarchical "deep research" (DR) agents that combine retrieval-augmented generation (RAG) over a local
- arXiv
- 2511.18303
- Published
- 2025-11-23
- Authors
- Rui Ding, Rodrigo Pires Ferreira, Yuxin Chen, Junhong Chen
AI summary
Overview
Research area: Large language model (LLM) agents for scientific discovery, specifically hierarchical "deep research" (DR) agents that combine retrieval-augmented generation (RAG) over a local document corpus with web search, applied to system-level materials and device discovery.
Technical level: Intermediate. The paper assumes familiarity with LLM agents, retrieval-augmented generation, and computational materials science terms (DFT, AIMD, OER, FET sensors), but the orchestration idea is explained conceptually.
Scope: The paper introduces and evaluates DToR (Deep Tree of Research), a locally deployable, tree-structured deep research agent, comparing 41 agents across 27 materials/device topics with LLM-as-judge scoring, pairwise duels, and five dry-lab simulation validations.
What This Paper Is About
Materials and device discovery problems at the "system" level (multi-layer stacks, core–shell–doped catalysts, nano-architected electrodes) involve combinatorial processing parameters scattered across heterogeneous, unstructured knowledge sources that existing machine learning surrogates and closed-source commercial agents do not handle well. The authors build a long-horizon, hierarchical deep research agent that runs locally, iterates between local-corpus RAG and web retrieval using LLM reasoners, and organizes its work as a tree of research branches that are adaptively expanded or pruned. The goal is to show that this orchestration produces higher-quality, source-grounded research reports than both a single-instance deep research agent and commercial deep research systems, at lower cost and with on-premises control.
Key Contributions
- A democratized, locally deployable framework for scientific deep research. Materials researchers can run open-source LLMs on their own hardware and exercise fine-grained control over cost and preferences, achieving performance that the authors report surpasses commercial state-of-the-art solutions in most cases.
- Hierarchical orchestration (DToR) that beats single-instance DR. A breadth-then-depth tree controller yields robust gains over a single deep research instance across LLM backbones (gpt-oss120B, gpt-oss20B, QwQ32B) and local-corpus budgets (local0/local100/local500), supported by full-factorial and targeted component ablations.
- An evaluation program combining three methods. Anchored rubric scores from LLM judges, repeated double-blind A/B duels, and dry-lab validation with domain simulations (including DFT and explicit-solvent AIMD) assess synthesis quality rather than retrieval exposure.
- An open-source release. The DToR framework source code is publicly available at https://github.com/ruiding-uchicago/DToR_deep_research.
Main Findings
- Best local agent ranks first overall. Across 27 topics and 41 agents, DToR_gpt-oss120B_local500 (DToR with the gpt-oss120B backbone and all 500 volumes of the local corpus) achieved an average rubric score of 8.57/10, ranking 1st. With local100 and local0 budgets it scored 8.53/10 and 8.37/10, both still top-3.
- Local agents beat commercial baselines. A single DR instance with gpt-oss120B reached 8.33/10 (local100) and 8.32/10 (local500), ranking 4th and 5th. DToR with gpt-oss20B produced 8.15/10, 8.12/10, and 8.08/10 for local500/local100/local0. All of these exceed ChatGPT-o4-mini-high (7.96/10), ChatGPT-o3 (7.95/10), ChatGPT-5-thinking (7.85/10), Gemini 2.5 Pro (7.81/10), Claude Opus 4 (7.78/10), Grok 3 (7.51/10), and Perplexity (7.51/10).
- A laptop-compatible backbone is competitive. DToR_QwQ32B_local500 attained 7.80/10, outperforming roughly half of the commercial systems.
- Gains concentrate on the hardest dimensions. Clarity and depth had the lowest means and largest dispersion. DToR improved on a single DR instance by an average of +0.69 (clarity) and +0.72 (depth). On novelty and applicability, DToR improved over Single_gpt-oss120B_local500 by +0.43 and +0.42 points on average, pushing above the leading commercial baseline. Relevance showed a ceiling effect and limited separability.
- Judge agreement was high. Average Pearson correlation p = 0.97 on raw scores and Spearman ρ = 0.97 on rank orders across the five critics.
- Orchestration is the dominant factor. In the Method × Backbone × Local-corpus budget factorial analysis, DToR consistently outperformed single-instance DR once a non-trivial local corpus was present (local100/local500); at local0 the gap narrowed and a strong backbone (gpt-oss120B) could close it. Enabling local RAG gave 7.21/10 and DToR mode 7.32/10 in the ablation.
- Every ablated component hurt. On DToR_gpt-oss120B_local500 across three representative topics, degrading tree depth (3 → 2 perspectives), reflection loops (3 → 0 iterations), and retrieval counts (5 → 1 papers per query) all reduced performance, with reflection proving most critical. Removing web search (local-only or LLM-only modes) also degraded quality. Driving the same orchestration with API models (gpt-5-mini/nano) maintained relative orderings but at materially higher cost.
- Duel results confirm the rubric. DToR agents showed a higher average win rate (58.6%) than Single DR agents (52.8%); DToR_gpt-oss120B_local500 reached a 79% mean win rate; gpt-oss-driven DToR and Single agents both exceeded 74%.
- Topic difficulty and discrimination are decoupled. Sensing/characterization tasks (in-situ TEM, printed FET arrays) were easiest; battery electrochemistry topics (LIB binders, fluoroether anion receptors) were hardest. Environmental sensing topics (microplastics, antibiotics detection) separated agents weakly (σ ≈ 0.7), whereas battery/materials synthesis challenges (LIB fluoroether receptor, ambient-pressure diamond growth) sharply distinguished strong from weak systems (σ > 1.1). LIB_Fluoroether_Anion_Receptor had both the highest difficulty (μ = 6.93) and the largest discrimination (σ = 1.24); Microplastics_Sensing_2D had the lowest discrimination (σ = 0.69) at mid-pack difficulty (μ = 7.19); Catalytic_Mineralization_of_Microplastics was relatively easy (μ = 7.49) yet discriminated well (σ = 1.03); Wastewater_Resource_Recovery_Cats was hard (μ = 7.04) but offered weak separation (σ = 0.93).
- Dry-lab validation supports actionability. Across 10 metrics on five application scenarios, there were four all-metrics > 100 events (every metric for a task exceeding its normalized baseline of 100): commercial once on CO2 sensor; local DR (non-gpt-oss) twice on CO2 sensor and battery binder; local DR (gpt-oss) once on PFAS degradation. The gpt-oss-based local DR agent led on 7 of 10 metrics by per-metric averages and achieved a higher overall aggregate score (98.7).
- Solvation reorders rankings. Explicit-solvent AIMD often reordered static-DFT rankings, revealing candidates that look promising in vacuum but lose ground when solvation and competition are included.
- A feasibility gap remains. The authors identify "kitchen-sink" failures, such as a PFAS sensing candidate (C3) stacking four mutually incompatible phases (hydrophobic fluorosilane, hydrophilic MXene, Ag nanoparticles, acidic PEDOT:PSS), and a PFAS degradation candidate (C5), a triple heterostructure (MXene/g-C3N4/Al2O3) that neglects interfacial charge-transfer resistance and fabrication complexity. They attribute these to "inverse-design hallucinations" arising from the absence of wet-lab feasibility priors, cost models, or external physics validators.
- Cost and runtime are tunable. On a single machine with serialized agent scheduling, DToR with gpt-oss120B averaged 19.6 h. A low-resource profile that disables local-corpus RAG finished in under 6 h, and a single-instance DR pipeline with gpt-oss120B typically completed within about 30 min. The best local model used 4.37 kWh per report on consumer devices. All runs used a workstation with 4 × NVIDIA RTX A5500 (Ubuntu).
Methodology in Plain English
The team starts from a "single DR instance": an evidence-first loop that generates a search query, retrieves from a local document corpus first, summarizes that local evidence, generates a complementary query for topical diversity, runs web research for both queries, updates a running summary, and reflects to propose a follow-up query until a maximum round count, then finalizes. Three choices distinguish it from common web-first agents: local-first retrieval, diversity-aware query generation, and robust input/output discipline to prevent stalls on locally hosted LLMs.
They then wrap that instance in a tree controller. Each DR instance becomes a "Research Node" in a branch-and-bound structure. A diversifier proposes several orthogonal perspectives, each seeding a branch with a budget (maximum depth, nodes per branch, total branches). For each active branch, a router runs the next pending node, an analyst decides EXPAND or PRUNE based on remaining depth, budget, and report quality, a knowledge gap explorer materializes new nodes targeted at identified gaps, and converged or depth-limited branches are synthesized into a perspective report. A final synthesizer reconciles cross-branch evidence, resolves conflicts, and outputs a provenance-rich consolidated report. Conceptually, this is a tree search over (query, evidence, summary) states rather than over purely symbolic "thoughts."
For evaluation, the authors used 27 expert-crafted topics spanning materials and devices (for example PFAS FET probes, CO2 sensing in 2D materials, Li/Na-selective membranes, CO2-reduction catalysts), 41 agents (11 commercial and 30 local, split evenly between single DR-instance and DToR mode), and five web-enabled judges (Claude 4 Opus (thinking), Gemini 2.5 Pro, Grok-3 (thinking), ChatGPT-o3, ChatGPT-o4-mini-high). Judges saw only the final report text — no agent identifiers, tool logs, or token metadata — and scored five equally weighted dimensions (relevance, depth, clarity, applicability, novelty) from 0–10 under a fixed "act as an experienced materials scientist" instruction, with positive and negative exemplars per dimension. Three repeats per topic–report–judge combination yielded 27 × 41 × 5 × 3 = 16,605 rubric judgments. Duels used three pools (commercial, local gpt-oss, local non-gpt-oss), ranked reports by rubric average, selected the top-3 in each pool, and ran all 36 pairwise duels per topic three times, yielding 2916 A/B trials and 14,580 individual judge decisions. Finally, five representative tasks — PFAS sensor probe, PFAS degradation catalysts, battery binder selection, oxygen evolution reaction (OER) catalyst stability, and CO2 sensor probe — were dry-lab validated: domain experts extracted candidates from anonymized best-local and best-commercial reports, built simulation environments, and scored 10 metrics per task with domain-standard protocols, against recognized baselines normalized to 100 (β-cyclodextrin for PFAS sensing, Ti4O7 (Magnéli) for PFAS degradation, PVDF for battery binders, IrO2 for OER catalysts, and g-C3N4 for CO2 sensors).
Why This Matters
The paper argues that device- and system-level discovery sits in a regime where governing laws for integration are not fully encoded in standard simulation packages, and where the combinatorial explosion of processing parameters is sparsely distributed across heterogeneous sources. Physics-aligned surrogates and domain LLMs operate largely at the molecular/crystal and small-assembly levels, while most agentic systems stop at intermediate depth and scope. DToR targets the gap with an open, on-premises framework that couples local domain RAG with gap-triggered web expansion inside a resource-bounded tree.
Real-world applications described or directly implied by the paper's tasks:
- Environmental sensing, including PFAS detection with 2D-material FET probes and CO2 sensing with candidate surfaces.
- Water treatment and resource recovery, including PFAS degradation catalysts and electrochemical recovery of Li+, PO4^3-, and NH4+ from complex wastewater.
- Energy storage, specifically advanced lithium-ion battery binder technologies for NCM811 cathodes.
- Electrocatalysis, specifically OER catalyst materials that resist corrosion in high-chloride, high-organic-load, or multi-ion wastewater streams.
Industry relevance: The framework runs on consumer-level hardware and open-source models, avoids commercial subscription fees, supports on-premises integration with proprietary local data and tools for privacy and controllability, and is tunable between quality (DToR plus local corpus) and cost/latency (single DR instance, no RAG). The authors position this as democratized access to advanced research agents, with branches parallelizable from lightweight laptops to multi-node clusters.
Future Directions
- Integrate domain-validator modules. The authors state that future AI4Science frameworks must add dedicated validators, such as synthesis-aware ReAct loops, to ground theoretical generation in experimental reality and reduce "inverse-design hallucinations."
- Add feasibility priors and cost models. Current agents lack intrinsic wet-lab feasibility knowledge, cost models, and external physics validators, which caused over-engineered proposals that ignored colloidal incompatibility, phase separation, acid-induced etching, interfacial charge-transfer resistance, and fabrication complexity.
- Improve the weakest rubric dimensions. Clarity and depth had the lowest means and largest dispersion across all agents, and even the best system was scored 8.57/10, leaving headroom.
- Scale and extend the evaluation. The study covers 27 topics and 41 agents with local RAG budgets of local0/local100/local500 and backbones gpt-oss120B, gpt-oss20B, and QwQ32B; the paper does not report evaluation over additional corpora sizes, model families, or topic domains beyond those listed, and the presented content is truncated before listing all topic prompts.
Target Audience
This paper is most useful to materials scientists and computational chemists exploring AI-assisted system-level design; machine learning researchers working on LLM agents, RAG, and hierarchical planning; and research computing or industry teams that need on-premises, cost-controllable research agents integrated with proprietary data and simulation tools. Readers primarily interested in wet-lab-validated synthesis recipes should note that the validation here is dry-lab (DFT and AIMD) rather than experimental.
Authors’ abstract
We present a long-horizon, hierarchical deep research (DR) agent designed for complex materials and device discovery problems that exceed the scope of existing Machine Learning (ML) surrogates and closed-source commercial agents. Our framework instantiates a locally deployable DR instance that integrates local retrieval-augmented generation with large language model reasoners, enhanced by a Deep Tree of Research (DToR) mechanism that adaptively expands and prunes research branches to maximize coverage, depth, and coherence. We systematically evaluate across 27 nanomaterials/device topics using a large language model (LLM)-as-judge rubric with five web-enabled state-of-the-art models as jurors. In addition, we conduct dry-lab validations on five representative tasks, where human experts use domain simulations (e.g., density functional theory, DFT) to verify whether DR-agent proposals are actionable. Results show that our DR agent produces reports with quality comparable to--and often exceeding--those of commercial systems (ChatGPT-5-thinking/o3/o4-mini-high Deep Research) at a substantially lower cost, while enabling on-prem integration with local data and tools.