Research
LABBench2: An Improved Benchmark for AI Systems Performing Biology Research
Overview Research area: AI for scientific discovery, specifically benchmark design for evaluating AI systems on practical biology research tasks. Technical level: Intermediate. The paper is readable w

- arXiv
- 2604.09554
- Published
- 2026-02-04
- Authors
- Jon M Laurent, Albert Bou, Michael Pieler, Conor Igoe, Alex Andonian, Siddharth Narayanan, James Braza, Alexandros Sanchez Vassopoulos, Jacob L Steenwyk, Blake Lash, Andrew D White, Samuel G Rodriques
AI summary
Overview
Research area: AI for scientific discovery, specifically benchmark design for evaluating AI systems on practical biology research tasks.
Technical level: Intermediate. The paper is readable without deep biology or machine learning expertise, though familiarity with language model benchmarks and agent tool use helps.
One-sentence scope: LABBench2 is a successor benchmark of more than 1,900 tasks that measures how well frontier AI models perform realistic biology research work across literature retrieval, data access, protocol troubleshooting, molecular biology, and experiment planning.
What This Paper Is About
The original Language Agent Biology Benchmark (LAB-Bench) measured AI abilities on practical biology tasks but used multiple-choice answering, in-line DNA sequences, and simplified contexts that no longer reflect what frontier models can do. LABBench2 rebuilds those task families in more realistic forms, such as open-response questions, PDF- and file-based inputs, and tasks that require retrieving the correct source paper, patent, clinical trial, or database entry. The goal is to measure whether models can actually perform useful scientific work, not just recall knowledge or reason in the abstract.
Key Contributions
- Expanded scope and realism: LABBench2 provides more than 1,900 tasks (1,912 according to Table 1) spanning literature understanding and retrieval, data access, protocol troubleshooting, molecular biology assistance, and experiment planning, with many tasks requiring retrieval or file-based inputs rather than in-line context.
- New literature evaluation variants: FigQA2 and TableQA2 each add three modes (image-provided, paper-provided PDF, and retrieval), and the benchmark adds three new task families — PatentQA, TrialQA, and SourceQuality — covering patents, clinical trials, and assessment of study quality.
- Baseline results for frontier models: The paper reports performance for current frontier models both without tools and with web search and code execution, showing a consistent difficulty increase over LAB-Bench.
- Public release: The dataset is released at huggingface.co/datasets/futurehouse/labbench2 and an evaluation harness at github.com/EdisonScientific/labbench2.
Main Findings
-
LABBench2 is substantially harder than LAB-Bench: Across all task families and models, accuracy drops consistently when moving from LAB-Bench to LABBench2, with model-specific differences ranging from −26% to −46%. This is attributed to higher-fidelity task framing such as open-response answering and retrieval- or file-dependent contexts.
-
Tool access helps unevenly: Tool augmentation (primarily web search and code execution) substantially benefits information retrieval tasks like LitQA3 and Patent/TrialQA. Accessing supplemental material also improves with tools, likely due to web search, but remains well below retrieval from main text — the paper speculates this is because supplemental material spans many file types including PDFs, Excel spreadsheets, CSVs, and figures.
-
Retrieval alone is not enough: FigQA2 in its default retrieve mode, SuppQA2, and DbQA2 remain bottlenecks. Even when models can reliably retrieve papers and web information, unlocking the information inside remains the limiting factor.
-
DbQA2 is among the hardest families: Database access tasks require navigating specialized scientific databases with nontrivial schemas, identifiers, and similarly named entities. DbQA2 remains one of the most challenging families in LABBench2 even with tools.
-
Standalone visual understanding is strong: For both FigQA2 and TableQA2, performance decreases moving from image-provided, to paper-provided, to retrieval modes. Models are impressive at reading figures and especially tables placed in front of them, but there is still a substantial gap in retrieving appropriate papers and locating figures within them.
-
Input modality drives molecular biology performance: In-line prompting generally gives marginally better results, though for models that support it, file access is nearly equivalent for SeqQA2. For Gemini 3 Pro and Opus 4.5 with tools, file-based CloningQA tasks actually appear easier than injected tasks, likely because of long input sequences often 3,000 base pairs or longer. Tool use substantially improves SeqQA2 and CloningQA, acting as a "great equalizer" across models. Retrieval-based variants perform poorly in all cases.
-
SeqQA2 performance is highly non-uniform: Subtasks requiring identification or manipulation of specific subsequences from larger context (such as primer design and amino acid identification) show the lowest performance across all models. More global operations (sequence complexity assessment, scoring sequence alignments) score higher. Some calculation-reliant tasks (GC content, enzyme kinetics) benefit especially from tool use.
-
A measurement caveat is reported: GPT 5.2 Pro does not accept appropriate file types with the Response API, which artificially deflates its file mode performance.
-
SourceQuality targets discernment rather than checklist compliance: This family is built from 100 open-access systematic reviews and asks models to justify, in open-ended form, why evidence-based medicine experts excluded a given study from a research question. The paper frames this as testing whether agents can surface the most epistemically salient exclusion rationale without being cued to a rubric.
Methodology in Plain English
The team kept the five broad categories of LAB-Bench but rebuilt the tasks so they resemble real research work rather than simplified quizzes.
Literature tasks (LitQA3, FigQA2, TableQA2, SuppQA2, PatentQA, TrialQA) were written by contracted domain experts holding or pursuing PhDs in biology or related fields, using a purpose-built web platform for task generation, review, and revision. Experts were instructed to ensure each question could only be answered using the intended source — and specifically not other parts of the same document — and that answers were not contradicted by any other source. Reviewers checked the same properties plus answer correctness.
FigQA2 and TableQA2 were each written in three modes: image only, full paper PDF, and retrieval with no context provided.
DbQA2's data sources were chosen by combining expert opinion with analysis: an LLM-based analysis of 1,000 recent biorXiv papers determined where input and output data was sourced or uploaded; top-ranked sources were manually curated and supplemented by internal expert input, yielding 43 sources.
SeqQA2 was built largely programmatically. Each of 20 subtask types was developed as a question template by an in-house biologist, then expanded to 20 variant questions using alternative input sequences. Where tasks have many valid answers (such as PCR primer design), custom verifier functions were written, for example performing in silico PCR with model-designed primers. Sequences can be delivered in-line, as files, or via retrieval.
ProtocolQA2 tasks were built by contracted domain experts who introduce an error into a protocol (from protocols.io, STAR Methods, or their own submissions — some unpublished) such that it causes an unambiguous negative result, then pose a scenario asking the model to identify an error causing the given result.
CloningQA requires complete end-to-end design of a molecular cloning protocol — DNA or enzyme reagents and all steps — spanning restriction-ligation, Gibson assembly, and Golden Gate methods. Answers are output in a simple domain-specific language that can be parsed and validated end-to-end in silico.
Evaluation compared base models (no tools) against models augmented with web search and code execution.
Why This Matters
Impact on research: LABBench2 argues that useful AI research systems must handle realistic workflows — finding the right source, navigating long heterogeneous documents, and extracting exact information — not just reasoning well. The paper positions itself as a continued de facto benchmark for AI scientific research capabilities.
Real-world applications:
- Therapeutic development, where retrieving and interpreting information from full-text patents and clinical trials is explicitly identified as high-value.
- Evidence synthesis and systematic review, where agents must identify why a study is inappropriate for a research question without relying on predefined checklists.
- Bioinformatics and data analysis, where agents navigate specialized scientific databases containing under-explored public data.
- Molecular biology and cloning workflows, where end-to-end protocol design and faithful sequence manipulation are required.
Industry relevance: The benchmark is designed for teams building foundation models, agents, and AI-driven research systems. The paper identifies specific engineering priorities — retrieval robustness with iterative search, better PDF and supplement parsing including multimodal artifacts, database navigation and API access, and purpose-built sequence manipulation tools. Prior near-saturation or superhuman results on LAB-Bench subcategories are cited, indicating that the field needs a harder target to keep measuring progress.
Future Directions
- Long-horizon task compositions spanning multiple task types, such as literature search to data access and analysis to protocol design to wet-lab experiments to result interpretation.
- Evaluation of ambiguous or non-deterministic outcomes, where there may be multiple valid plans or partially correct approaches.
- Focused evaluations in specific scientific sub-domains, including high-impact areas like drug development, to isolate capabilities more cleanly.
- More reliable evaluation schemes and representative tasks, since LABBench2 still simplifies wet-lab execution, ambiguous goals, and long-horizon iteration across days or weeks, and grading depends on the reliability of verifiers.
Target Audience
AI researchers and engineers building agents for scientific work; benchmark developers interested in realism and grading design; biology and bioinformatics practitioners evaluating whether AI tools can support real research workflows; and research organizations or labs assessing where current models still fail. Readers seeking exact per-model accuracy figures should note that the summary text does not report them individually — the paper presents them primarily through Figure 1, Figure 2, Figure 3, and Figure 4, with only the −26% to −46% difficulty range stated numerically.
Authors’ abstract
Optimism for accelerating scientific discovery with AI continues to grow. Current applications of AI in scientific research range from training dedicated foundation models on scientific data to agentic autonomous hypothesis generation systems to AI-driven autonomous labs. The need to measure progress of AI systems in scientific domains correspondingly must not only accelerate, but increasingly shift focus to more real-world capabilities. Beyond rote knowledge and even just reasoning to actually measuring the ability to perform meaningful work. Prior work introduced the Language Agent Biology Benchmark LAB-Bench as an initial attempt at measuring these abilities. Here we introduce an evolution of that benchmark, LABBench2, for measuring real-world capabilities of AI systems performing useful scientific tasks. LABBench2 comprises nearly 1,900 tasks and is, for the most part, a continuation of LAB-Bench, measuring similar capabilities but in more realistic contexts. We evaluate performance of current frontier models, and show that while abilities measured by LAB-Bench and LABBench2 have improved substantially, LABBench2 provides a meaningful jump in difficulty (model-specific accuracy differences range from -26% to -46% across subtasks) and underscores continued room for performance improvement. LABBench2 continues the legacy of LAB-Bench as a de facto benchmark for AI scientific research capabilities and we hope that it continues to help advance development of AI tools for these core research functions. To facilitate community use and development, we provide the task dataset at https://huggingface.co/datasets/futurehouse/labbench2 and a public eval harness at https://github.com/EdisonScientific/labbench2.