Research
InvestigationWorlds: An Agentic Environment for Legal Investigation
InvestigationWorlds: An Agentic Environment for Legal Investigation Overview Research area: Artificial intelligence / agentic evaluation benchmarks, with an application domain in U.S. civil litigation

- arXiv
- 2610.04129
- Published
- 2026-10-02
- Authors
- Albert Yu Sun, Andrew Benard, Sil Hamilton, Anna Teresita A. Marcelo, Yong Jae Kim, Carl-Leander Henneking, Rundong Hu, Yuhong Wang, David Mimno, Bishan Yang, Igor Labutov
AI summary
InvestigationWorlds: An Agentic Environment for Legal InvestigationOverview
Research area: Artificial intelligence / agentic evaluation benchmarks, with an application domain in U.S. civil litigation (legal investigation and e-discovery).
Technical level: Intermediate. The agentic evaluation setup, scoring, and ablations are straightforward to follow; familiarity with retrieval-augmented agents and legal-discovery terminology helps but is not required.
Scope (one sentence): The paper introduces InvestigationWorlds, an agentic environment built from 100 real U.S. federal summary judgment cases via PACER and augmented with attorney-validated synthetic, role-tagged documents, in which agents must investigate a corpus and commit to the court-adopted factual hypothesis rather than competing alternatives.
What This Paper Is About
Investigative reasoning often requires assembling a coherent factual account from large collections of documents that can support several competing interpretations — the situation attorneys face during discovery, which the paper notes accounts for an estimated 90% of litigation costs and time (citing Girard and Espinosa, 2010; Withers, 2000). Existing benchmarks do not measure this capability: they test legal question answering, precedent retrieval, hallucination in commercial research tools, and bar-exam-style reasoning, but they assume either a single correct answer or a well-posed query, and do not create settings where the same evidence supports multiple coherent but incompatible conclusions. The paper's goal is to build an environment with judicially validated ground truth where an agent must construct its own account of a case before seeing a menu of candidate hypotheses.
Key Contributions
-
Investigation environments grounded in real court cases. InvestigationWorlds is described as the first agentic environment with multiple adversarial interpretations grounded in real cases. The authors introduce an attorney-validated generation procedure that takes a real case record and produces a corpus admitting multiple plausible factual findings, with each synthetic document tagged by its evidentiary role. The court's opinion supplies the ground truth used to measure whether an agent predicts the court-adopted disposition. They release the full PACER document corpus for all 100 cases; because documents were retrieved with the RECAP browser extension (Free Law Project; Yu and Schultze, 2011), every purchased document has been contributed to the Free Law Project's public RECAP Archive on CourtListener.
-
A reproducible pipeline. The authors release (i) a generation pipeline including all prompts and the attorney-validated role taxonomy, and (ii) a portion of the generated synthetic documents, anonymized for ethical and privacy reasons. The dataset and synthetic data generation code are released at https://huggingface.co/datasets/investigation-worlds/investigation-worlds.
-
A two-turn investigation protocol and a case-solved metric. The environment is defined as a triple (corpus, entities, hypothesis DAG). The hypothesis DAG contains the court-adopted hypothesis, the court-rejected hypothesis, and synthetic alternative hypotheses. The evaluation separates investigation (Turn 1, tools enabled) from commitment (Turn 2, tools disabled), and scores "case-solved" as 1 only when the model believes the court-adopted hypothesis and rejects every alternative.
-
A role taxonomy for synthetic evidentiary documents. Five roles are defined and attorney-validated: Signal (embeds ground facts in a realistic discovery artifact), Noise (shares parties, contracts, and date windows with signal documents but embeds no ground facts), Chaff (same business world but no shared parties, dates, or dispute topics), Pro-alternative signal (embeds a fact tilting toward an alternative hypothesis while remaining individually compatible with every ground fact), and Alternative-specific noise (consistent with an alternative's parties and timeline but embeds none of its facts).
Main Findings
-
The best-performing model solves only one-third of cases. On the 100-case cohort, Claude Haiku 4.5 achieves the highest case-solved rate (33.3%), followed by Claude Sonnet 4.6 (29.9%). All six evaluated models exceed the uniform-random verdict baseline of 9.9%, but only these two exceed the stronger baseline that believes one hypothesis at random and rejects the rest (25.0%). Their gains over that baseline are 8.3 and 4.9 percentage points, and neither improvement is statistically significant.
-
Failure is not a retrieval failure. Five of the six models achieve 41.9%–52.6% accuracy on their worse-performing class, exceeding the maximum worst-class accuracy of 33.3% among the evaluated naive policies; GPT-5.4-mini is the exception at 20.0%. Separately, models open only 0.7–10% of the corpus on average yet recover signal documents at substantially higher rates than uniform retrieval, and none of the strategy differences (coverage and signal recall) predict accuracy. The authors conclude retrieval volume is not the bottleneck.
-
Commitment behavior changes the ranking. On the 87 cases with complete outputs from both Anthropic models, Haiku 4.5 commits on 80.7% of verdicts versus 73.3% for Sonnet 4.6, while Sonnet is more accurate among committed verdicts (74.1% versus 70.8%). Haiku and Sonnet believe the court-adopted hypothesis in 45 and 46 cases respectively, but also believe an incorrect alternative in 16 and 21 of those cases; Haiku solves 29/87 cases (33.3%) versus 25/87 (28.7%) for Sonnet.
-
Pro-alternative documents pull agents toward wrong hypotheses. In a five-level cumulative ablation on 50 cases starting from a merits-only baseline, adding 20 pro-alternative documents reduced case-solved rates for five of six models, with declines ranging from 4 to 22 percentage points; paired 95% bootstrap confidence intervals exclude zero for Haiku 4.5 (−22 points) and GPT-5.4-mini (−8 points). GPT-5.5 increased by 2 points. Additional noise and chaff produced no consistent further decline.
-
Belief in alternatives rises sharply with adversarial evidence. Across six models, 50 cases, and two alternatives per case, 506 paired verdicts with parsed answers in both conditions were retained (94 pairs excluded for missing or unparsed output). Belief in reference-incorrect alternatives increased from 3/506 (0.6%) to 115/506 (22.7%): 55 answers changed from "do not believe," 59 from "do not know," and one remained "believe."
-
Investigative strategies vary widely but do not explain accuracy. Sonnet 4.6 anchors the high-effort corner (7.4% coverage, 24.7% signal recall) and DeepSeek V4 Flash reads more documents on average (10.0% coverage, 24.2% signal recall). Haiku, Qwen-3.6-35B, and GPT-5.5 form a middle cluster at 3–6% coverage and 15–19% signal recall, while GPT-5.4 mini sits near the origin (0.7% coverage, 3.1% signal recall).
-
Evaluator rankings are robust to the generator. Rebuilding 30 held-out cases end-to-end with OpenAI models in place of Anthropic models throughout the pipeline preserved evaluator rankings up to ties (Spearman ρ = 0.94), and no pair of evaluators reversed its strict ordering. Haiku 4.5 and DeepSeek v4 tie at 26.7% on OpenAI-generated corpora; on Anthropic-generated corpora, DeepSeek v4 and GPT-5.5 tie at 10.0%, while Qwen 3.6 35B and GPT-5.4-mini tie at 6.7%. DeepSeek v4 showed the largest change, from 10.0% to 26.7% (paired difference +16.7 percentage points; 95% CI [0.0, 33.3]).
-
The environments are grounded and human-verifiable. An attorney who had not audited the cases investigated 10 randomly selected cases across the five federal categories with the court's opinion and adopted label withheld, reading 459 pages over 12 hours; their factual conclusions agreed with the court-adopted outcome in all 10 cases. Across the 55-case deep-review subset, pass rates exceeded 94% for party hypotheses, alternative hypotheses, and chaff documents, at an 81%+ agreement score on the 10-case agreement subset. Pooled across all stages (2,009 items, 55 cases), pass was 81.3%, borderline 10.9%, fail 7.8%, with 78.2% pairwise agreement.
Methodology in Plain English
Sourcing ground truth from real cases. The authors exploit the summary judgment motion, in which both sides must lay out every material fact they rely on, attach evidentiary exhibits, and either accept or dispute the opposing party's facts one by one; the judge then issues a written opinion adopting specific factual findings traceable to those exhibits. This yields a complete evidentiary corpus paired with a judicially validated answer.
Case selection. The filter selects cases with fully decided summary judgment motions where the statements of facts, counter-statements of facts, evidentiary exhibits, and the court's opinion are available and unsealed. The cohort spans 100 cases from 28 states, the District of Columbia, and Puerto Rico across 41 federal judicial districts in all 12 regional circuits, balanced 20-per-category across 5 federal case types, with 42,283 pages in the archived downloaded PACER PDFs (median 290 pages per case, maximum 3,274). Only "merits exhibits" are exposed to the agent: procedural filings and filings that would reveal the answer (such as expert opinions) are discarded.
Extraction. An LLM (Claude Sonnet 4.6) reads the opinion to extract the court's actual factual findings, the legal questions, and supporting text passages, while the pipeline parses the statements of facts and the opposing counter-statement. The same model extracts and disambiguates entities — people, organizations, instruments, artifacts — so that references to the same actor collapse to one identity, and annotates each exhibit with document type, relevant date, and mentioned entities.
Augmentation. Because a real merits record is much smaller than a real investigation corpus, an LLM (Claude Sonnet 4.6 by default) generates two alternative hypotheses per case, each paired with the reasoning that makes it superficially plausible and the facts that refute it. The model plans documents in coherent series (for example, a purchase order, its matching invoice, and the corresponding wire-transfer record) and assigns each specification a role, a target fact or hypothesis, and a rendering format (pdf, docx, xlsx, pptx, or plain text). Sub-agents render each document; a judge model (Claude Haiku 4.5) reviews each finished document against its specification, and failures are rejected and regenerated up to three times.
Validation. A litigation attorney pre-screened the case set before generation. Three practicing attorneys with 30+ years of combined experience then audited pipeline outputs, completing approximately 100 hours of expert review. All 100 cases received a hypothesis-fidelity audit; 55 cases received a stage-by-stage review; within those, an agreement subset of 10 cases drawn proportionally across the five federal categories was independently reviewed by all three attorneys.
Evaluation protocol. The hypothesis set per case is {h_p, h_d, h_1, h_2} — two party hypotheses from the plaintiff's and defense's briefs and two pipeline-generated alternatives. The court-adopted hypothesis h* is identified from the opinion; the other party hypothesis becomes the court-rejected hypothesis h†. Ground truth is "believe" for h* and "do not believe" for the other three. In Turn 1 the agent receives the case type and corpus, tools enabled, and produces an investigation summary with citations. In Turn 2 the agent receives only its own summary and a shuffled list of hypotheses labeled H1, H2, …, with tools disabled, and returns "believe," "do not believe," or "do not know" for each. "Do not know" on an alternative counts as not believing it. Six frontier agents were evaluated: Claude Sonnet 4.6 and Claude Haiku 4.5 (Claude Agent SDK), GPT 5.5 and GPT 5.4 mini (OpenAI Agents SDK), and DeepSeek V4 Flash and Qwen 3.6 35B (vLLM-backed mini-swe-agent harness). Both harnesses expose the same four file-system operations (read, glob, grep, shell), with a 30-tool-call budget and 25 agent turns in Turn 1. The full project consumed approximately $3,300 of API spend, with per-case investigation costs from $2.06 (Sonnet 4.6) to $0.07 (GPT 5.4 mini).
Why This Matters
Impact on research. The paper argues that no existing benchmark measures investigation — navigating documents with minimal guidance to construct competing factual hypotheses — because existing evaluations assume a single correct answer or a well-posed query. By pairing real records with judicially adopted ground truth and adversarially constructed alternatives typed by evidentiary role against a hypothesis DAG, it isolates a failure mode distinct from retrieval: agents can retrieve relevant evidence and still commit to the wrong account. It also introduces a template for creating investigation environments wherever institutional decisions produce reliable ground truth.
Real-world applications (implied by the setting):
- Legal discovery and e-discovery review, where teams must read across large productions to determine which narrative the evidence supports.
- Litigation support and case assessment, including testing whether a tool's factual conclusions match what a court ultimately adopted.
- Evaluation and auditing of commercial legal AI tools, extending beyond question answering and hallucination checks to narrative reconstruction.
- Extension to other adjudicative settings the authors name: state-court records, administrative adjudications, and arbitration awards.
Industry relevance. The paper reports that discovery accounts for an estimated 90% of litigation costs and time (citing Girard and Espinosa, 2010; Withers, 2000), and it evaluates six frontier agents from Anthropic, OpenAI, DeepSeek, and Qwen. The finding that the strongest case-solved rate is 33.3% — and that gains over a naive one-hypothesis baseline are within noise — is directly relevant to anyone deploying legal agents. The authors further report a practical tradeoff: a model that abstains more often (Sonnet 4.6) can rank below a model that commits more often (Haiku 4.5), which matters when choosing models for decision-support workflows.
Future Directions
- Broadening jurisdiction and case type. The environments come only from U.S. federal
Authors’ abstract
We introduce InvestigationWorlds, an agentic environment for legal investigation. We build on an underused artifact of U.S. civil litigation: the summary judgment motion. This motion relies upon a record composed of real evidence exhibits, and results in a court-adopted hypothesis that is treated as ground truth for the purposes of deciding the motion. Each environment is built from a real U.S. Federal Court case retrieved from Public Access to Court Electronic Records (PACER) and augmented by an attorney-validated generation pipeline that synthesizes role-tagged documents around the original record. The resulting corpus admits multiple coherent factual readings, only one of which matches the court-adopted hypothesis. Evaluating on 100 cases, we find agents often commit to incorrect hypotheses despite retrieving relevant evidence, struggling to distinguish the court-adopted hypothesis from alternative hypotheses.