Research
LongCat-DeepResearch Technical Report
LongCat-DeepResearch Technical Report Overview Research area: Autonomous "deep research" agents — language-model systems that investigate open-ended questions, gather evidence from external sources, a

- arXiv
- 2609.36071
- Published
- 2026-09-28
- Authors
- Meituan LongCat Team, He Zhu, Yue Xu, Wanli Wu, Haolin Ren, Yuxin Bian, Jiarui Zhao, Rongzhi Zhang, Quanchi Weng, Jinghao Cui, Yu Fan, Yuhan Liu, Yunhu Ye, Jiyuan Ren, Fengcheng Yuan, Zhao Yang, Jiacheng Zhang, Yuchuan Dai, Ruixuan Xiao, Haozhe Sun, Xiangyuan Liu, Cheng Sun, Yao Du, Yiming Hao, Hongbo Guo, Shuo He, Lei Wang, Xunliang Cai, Yan Chen, Fan Yang, Lingchuan Liu
AI summary
LongCat-DeepResearch Technical ReportOverview
Research area: Autonomous "deep research" agents — language-model systems that investigate open-ended questions, gather evidence from external sources, and synthesize long-form, citation-bearing reports. The paper sits at the intersection of multi-agent orchestration, context management, and training-data construction for research-capable models.
Technical level: Intermediate. The report is a systems paper rather than a purely algorithmic one: the individual ideas (planning, retrieval, sectioned writing, editorial revision) are easy to grasp, but the evaluation protocol, benchmark-specific rubrics, and ablation structure require some familiarity with how deep-research systems are scored.
Scope: The paper describes LongCat-DeepResearch, a harness centered on an executable specification called ResearchSpec, combined with an enhanced LongCat model, and reports its scores on DeepResearchBench, DeepResearchBench II, ResearchRubrics, and an in-house benchmark alongside component ablations and design analyses.
What This Paper Is About
Open-ended deep research cannot be fully specified in advance: as an analysis develops, new questions surface, claims need more verification, and evidence requirements change. The paper asks how a research agent can keep a shared research agenda while still carrying out the detailed, separate investigation each section requires. Its goal is a workflow that keeps global coordination compact — concentrated in a small specification rather than an ever-growing draft — while letting each section's research happen in its own independent context, followed by targeted rather than whole-report revision.
Key Contributions
-
An executable planning artifact (ResearchSpec). Multiple planning agents search and read external sources, propose candidate plans, and their proposals are consolidated and refined by a Planning Judge, a Critic, and a Reviser. ResearchSpec records each section's scope, research questions, required entities or cases, and provisional source leads, with unique hierarchical IDs such as S1.2 and validators checking ID uniqueness, parent relationships, consecutive ordering, and required fields.
-
A decoupled research-then-edit harness. Each Researcher receives the original query, the complete ResearchSpec, and one section assignment, gathering evidence and writing a complete citation-bearing section in its own context. Sections are then mechanically assembled in ResearchSpec order, a Global Editor assigns ownership of repeated material and specifies changes, and Local Editors apply those directives to assigned sections.
-
Stage-level interfaces for research data construction. Planning, research, and editing each define a mapping that can be used to synthesize targeted data: research tasks with evidence-backed requirements, task-specific rubrics (including implicit requirements and negative conditions), and trajectories recording candidate and final ResearchSpecs, tool calls, sections, assembled drafts, and editorial decisions. These are used in the mid-training and post-training of LongCat's general-purpose models.
-
A multi-benchmark evaluation with design analyses. Results on three public benchmarks plus an in-house benchmark, plus component ablations, ResearchSpec refinement study, Editor scaling study, and model–harness configuration comparisons.
Main Findings
-
Public benchmark scores: LongCat-DeepResearch obtains 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics, exceeding the strongest of the three compared web products by +0.30, +3.17, and +5.62 points respectively. The paper states these point estimates use the listed scored cohorts and do not establish statistical significance.
-
Where the DeepResearchBench gains come from: LongCat-DeepResearch has the highest comprehensiveness (56.21), insight/depth (56.48), and instruction following (55.18) scores, with an overall of 55.25 versus ChatGPT-DeepResearch's 54.95. Readability is weaker, at 49.96, below ChatGPT-DeepResearch's 51.51 and Claude-DeepResearch's 51.35.
-
DeepResearchBench II strength in analysis: Overall 51.35, exceeding Claude-DeepResearch by 3.17 points. It leads in information recall (46.42) and analysis (60.69, versus 52.26 for the next-highest system), while its presentation score of 83.13 trails ChatGPT-DeepResearch's 86.98.
-
ResearchRubrics strengths and remaining gaps: Overall 79.83, exceeding ChatGPT-DeepResearch by 5.62 points. It leads in explicit requirements (86.97), implicit requirements (79.52), and synthesis (78.92), but instruction following (66.27), citation quality (65.00), and communication (60.00) each remain below at least one compared system.
-
In-house benchmark: LongCat-DeepResearch ranks second among four systems with an overall score of 76.04, above Claude-DeepResearch at 61.42 and Gemini-DeepResearch at 42.49, and 0.55 points below ChatGPT-DeepResearch's 76.59. It has the highest Content score (87.11) and Analysis score (70.74), with Presentation at 52.99.
-
Component ablations: On DeepResearchBench II and ResearchRubrics, the Full pipeline scores 48.68 and 79.14 (average 63.91). Removing the Planning Judge and Critic and running one Writer then Reviser gives 44.65 and 74.46 (average 59.56); using one whole-report Researcher gives 45.64 and 78.49 (average 62.07); using no Editor gives 48.60 and 78.00 (average 63.30).
-
ResearchSpec aggregation helps; further refinement has mixed effects: On a fixed ResearchRubrics development subset, planned coverage rises from 53.76 (single Writer, W0) to 60.47 (Judge-integrated, J0), 63.14 (one Critic/Reviser cycle, R1), and 62.48 (two cycles, R2), while the full-rubric report Total moves 81.67 → 85.13 → 84.34 → 84.89. Plan aggregation improves final-report quality; refinement further improves planned coverage.
-
Editor scaling improves readability preference: Across enhanced Editor rounds on fixed development subsets, the average native score moves from 70.89 (no Editor, A0) to 70.85 (original Editor, A1), 70.17 (B1), 71.28 (B2), and 72.19 (B3), while average automatic readability preference (50 denotes parity) moves 50.00 (A1), 50.42 (B1), 53.33 (B2), 53.96 (B3). The paper notes benchmark-specific trends differ.
-
Harness change matters as much as the model: With the previous LongCat release (LongCat-2.0), switching from ReAct/direct report to the current harness raises the three-benchmark unweighted average from 47.04 to 58.05, and switching to the current model with the current harness raises it to 62.14.
-
Reported limitations of the evidence: FACT results on DeepResearchBench are not reported; the paper notes FACT measures statement–page support rather than source authority or claim truth, and that the bounded searchability check does not certify exhaustive web answerability. System-level comparisons do not isolate a model or harness component.
Methodology in Plain English
The system replaces early iteration on a full report with iteration on a compact plan.
-
Plan first, but explore while planning. Independent Planning Writers briefly search and read external sources before proposing a report structure, so the plan is grounded in evidence beyond the model's prior knowledge. A Planning Judge consolidates the candidate specifications, a Critic searches for missing questions and cases, and a Reviser folds that feedback in. The refined output is ResearchSpec: a per-section list of scope, research questions, required entities or cases, and source leads. Validators check structure; invalid plans are repaired or revert to a validated candidate, and unresolved errors stop dispatch.
-
Research each section in its own context. Every Researcher gets the original query, the full ResearchSpec, and one assignment. Researchers do not depend on previously written sections, run concurrently, and write their own citation-bearing sections from the evidence they collected — so local evidence goes straight into writing rather than being compressed for a separate writer. ResearchSpec is fixed during this stage.
-
Assemble, then edit selectively. Sections are assembled in ResearchSpec order, retaining text and citations even when content overlaps. A Global Editor reads the whole draft, assigns ownership of repeated material, and specifies changes; Local Editors then apply directives to their assigned sections, using the full draft and specification as context. Because the output scope is local, one editorial call does not regenerate the whole document.
-
Turn the interfaces into training data. A target profile defines language, topic, breadth, and intended report. Questions and rubrics are built from either an independently licensed review article or frozen multi-source briefs, and factual criteria must have supporting evidence while analytical criteria specify warranted comparisons or inferences. A bounded searchability check labels atomic recall criteria as Keep, Revise, or Drop, and a gate passes only if all retained criteria receive Keep. Deterministic checks cover duplicate rubrics, answer leakage, excluded-source leakage, and temporal scope. Accepted queries then drive teacher executions of the harness, and trajectories are filtered for role identity, tool-call/response closure, and intermediate-artifact validity while preserving complete, partial, and failed attempts.
-
Evaluate with each benchmark's official setup. DeepResearchBench RACE and DeepResearchBench II are judged by GPT-5.5 (medium), and ResearchRubrics by Gemini 2.5 Pro. The three product comparisons use each provider's own Deep Research system through its official client: Gemini 3.7 Flash, GPT-5.6 Sol (xhigh), and Claude Opus 5 (xhigh). The in-house benchmark uses an automatic evaluator.
Why This Matters
Impact on research. The paper makes a concrete design argument about where iteration should happen in deep-research systems: not in an expanding draft, but in a compact, inspectable specification of what the report needs to establish. It also treats planning and editing as separate, checkable interfaces, which makes them usable as units for data synthesis and rubric-based evaluation — potentially changing how research-capable models are trained, not just how they are prompted. The reported gains when the harness changes while the model is held fixed (47.04 to 58.05 average) suggest workflow design is a first-order factor rather than an implementation detail.
Real-world applications implied by the work:
- Producing long, citation-bearing literature reviews and technical reports where claims must be tied to sources.
- Investment and finance research: the paper's case study works through an AI-enhanced portfolio management question covering classical theories, AI/ML applications, multi-criteria decision making, and portfolio optimization and rebalancing through April 2024.
- Multi-section analytical deliverables that require reconciling overlapping evidence across separately researched parts, such as market or policy briefs.
- Generating research-oriented training data — questions, task-specific rubrics, and agent trajectories — for improving general-purpose models.
Industry relevance. The comparison set is made of deployed commercial deep-research products, and the paper reports results within a few points of the strongest one (76.04 versus 76.59 on the in-house benchmark, while leading on all three public benchmarks). The harness is described as separable from the model, which is relevant to teams that must decide whether to invest in a better model or a better research workflow. The paper's own caveats matter commercially too: coverage differs across systems, costs can be substantial when the full draft is repeatedly fed back into editors, and the comparisons are system-level, so they do not attribute gains to any single component.
Future Directions
- Reopening the plan after research. The paper explicitly states that revising the specification during section research, or a later research round that reopens the plan, is an extension beyond the current pipeline.
- Making planning refinement reliably beneficial. Additional planning candidates and refinement rounds cost more computation, and the paper reports their benefit is not guaranteed and that refinement has mixed effects on report quality even as planned coverage improves.
- Preserving useful detail through editing. The paper notes that research and editing can omit useful details, that repeated full-draft inputs can still incur substantial cost, and cites direct evidence elsewhere that revision can damage earlier coverage or citation quality; the Editor scaling study shows readability preference improving while native benchmark trends differ.
- Isolating what actually drives the scores. The system-level scores do not isolate the contribution of the data pipeline, any single data source, or any one training stage, and FACT results are not reported.
- Establishing statistical confidence and fair coverage. The public-benchmark differences are presented as point estimates without statistical significance, and the paper notes scored coverage varies across systems and that detailed coverage and execution records are in the appendix.
Target Audience
Researchers and engineers working on autonomous research agents, multi-agent orchestration, or long-form generation with citations; teams building or evaluating commercial deep-research products who need to reason about harness design versus model capability; and practitioners constructing synthetic tasks, rubrics, and agent trajectories for training. The paper is most useful to readers who care about workflow architecture and benchmark-level evidence rather than a new single model or a formal theoretical result, and readers should approach the score comparisons alongside the paper's own caveats about significance, coverage, and component attribution.
Authors’ abstract
We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports. The workflow separates global planning from detailed investigation and coordinates revision at the section level. Multiple planning agents first explore external sources and refine an actionable research plan, termed ResearchSpec. Research agents then investigate and draft their assigned sections in parallel, gathering additional evidence in separate contexts as their analyses develop. Once the sections are assembled, global review guides targeted local revisions, reducing reliance on repeated full-report rewriting. This workflow also supports the construction of research tasks and trajectories for the mid-training and post-training of LongCat's general-purpose models. LongCat-DeepResearch achieves 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics. On an in-house benchmark, it scores 76.04, ranking second among four compared systems. Development-set analyses show benefits from combining planning perspectives, while further planning refinement has mixed effects. Additional editing improves average automatic readability preference across two benchmarks, with different trends on each.