Skip to content
AI.info

Research

Securing Dual-Use Pathogen Data of Concern

Overview Research area: AI governance and biosecurity, specifically the control of dual-use pathogen data used to train biological AI models. Technical level: Beginner-Friendly. The paper is a policy

arXiv
2602.08061
Published
2026-02-08
Authors
Doni Bloomfield, Allison Berke, Moritz S. Hanke, Aaron Maiwald, James R. M. Black, Toby Webster, Tina Hernandez-Boussard, Oliver M. Crook, Jassi Pannu

AI summary

Overview

  • Research area: AI governance and biosecurity, specifically the control of dual-use pathogen data used to train biological AI models.
  • Technical level: Beginner-Friendly. The paper is a policy and taxonomy proposal written in largely non-mathematical prose; readers need no machine-learning background, though familiarity with biosecurity policy helps.
  • Scope in one sentence: The paper proposes a five-tier Biosecurity Data Level (BDL) framework for classifying pathogen data by misuse risk, pairs each tier with escalating technical controls, and outlines a governance body to implement the system.

What This Paper Is About

Biological AI models are trained on large volumes of pathogen-related data, and the type of data used strongly influences what the resulting model can do, including capabilities that could aid bioweapons development. Most AI safety governance focuses on model training, model weights, or model outputs, while the question of which datasets should be freely available upstream is comparatively neglected. This paper proposes a way to categorize the narrow subset of newly created pathogen data that could confer concerning capabilities on AI models, and to restrict access to it without shutting down the broader norm of open biological data sharing.

Key Contributions

  1. A five-tier Biosecurity Data Level (BDL) taxonomy, ranging from BDL-0 to BDL-4, that sorts pathogen data types by their expected contribution to AI capabilities of concern. The framework is explicitly modeled in naming only on the Biosafety Level (BSL) system; the authors state there is no other relationship between the two systems and no correspondence between a BDL level and experiments performed at a given BSL level.
  2. A matched set of technical data-access controls, presented in Table 1, that escalate with risk: optional monitoring at BDL-0, managed access at BDL-1, identity accreditation at BDL-2, use approval at BDL-3, and pre-screening at BDL-4, with specific security methods assigned to each tier.
  3. An operational specification for Trusted Research Environments (TREs) tailored to biological AI training data, organized around the Five Safes framework (Safe People, Projects, Data, Settings, Outputs) and adapted for biosecurity rather than only privacy.
  4. A proposed governance framework, centered on four principles drawn from past U.S. research oversight, including the creation of an independent Pathogen Data Board to write and revise classification rules.

Main Findings

  • Viral data exclusion degrades model performance on virus tasks. Excluding viral proteins from the training data of the generative protein model ESM3 "noticeably reduced its performance on benchmarks related to viruses" (Hayes et al., 2025). Excluding genomes of viruses infecting eukaryotes from the training data of the genomic language model Evo 2 "led to poor performance on fitness-prediction tasks related to the genomes of viruses infecting humans" (Brixi et al., 2025).
  • Fine-tuning can restore capabilities that pretraining excluded. Work by a subset of the paper's authors shows that fine-tuning Evo 2 on sequences from human-infecting viruses, which were originally excluded from the Evo 2 training set, reduced the model's perplexity in predicting withheld human-infecting viral sequences compared with the pretrained model. The fine-tuned model also outperformed the pretrained model at predicting SARS-CoV-2 immune escape variants (Black et al., 2025).
  • The BDL tiers are defined by the properties data would let a model learn. BDL-0 covers all data types not named in the higher tiers. BDL-1 covers data enabling a model to learn general patterns of eukaryote-infecting viral genome and protein composition. BDL-2 covers properties of viruses from families likely capable of causing pandemics in non-human animals, such as zoonotic crossover, environmental stability, host range, susceptibility of host populations, and evasion of diagnostics or detection. BDL-3 covers transmissibility, virulence, immune evasion, and resistance to medical countermeasures for viruses likely to infect humans. BDL-4 covers enhanced transmissibility, enhanced virulence, or enhanced immune evasion relative to known wild-types for viral variants from families likely to cause pandemics in humans.
  • Control methods are assigned by tier. BDL-1 pairs managed access with provenance and audit logs, anomaly detection, API rate limiting, risk scoring, and intent logging. BDL-2 adds hardware keys, enhanced cross-institutional anomaly detection, and watermarking. BDL-3 adds TRE integration with confidential computing, behavioral biometrics, sophisticated honeytokens, and real-time risk scoring. BDL-4 adds TRE integration with provenance recording, facial recognition, and cross-institutional federated learning.
  • The taxonomy is presented as provisional. The authors assume that functional data, such as deep mutational scanning or virus-host protein-protein interaction data, is particularly relevant to enabling capabilities of concern, because cellular and organismal complexity make it hard to predict functional outcomes from genetic information alone. They note that some models predicting viral fitness rely only on sequencing data (e.g., King et al., 2025), and state that if biological foundation models generalize well from large volumes of limited data types such as sequence data, the BDL framework would require revision.
  • Four governance principles are drawn from prior U.S. oversight. Rules should be readily applicable with little room for subjective local judgment; oversight should be enforced by a neutral arbiter rather than through end-to-end self-regulation; researchers should get a guaranteed timeline for classification decisions and a clear appeal route; and rules should bind all scientists, not only government grantees.
  • The framework is aimed at newly created data. The proposal targets newly created dual-use pathogen data of concern rather than preexisting data, and recommends removing BDL restrictions on data about pathogens causing a significant and ongoing disease outbreak.

Methodology in Plain English

The authors did not run new experiments. They assembled existing empirical evidence about how removing or adding viral data changes model behavior, then used that evidence to argue that data is a leverage point for biosecurity. From there, they constructed a classification scheme by reasoning about what a model would need to learn, and what data types would teach it, in order to acquire capabilities such as predicting or generating more transmissible or immune-evasive pathogens. Each tier was then matched to security tools borrowed from cybersecurity and from the health-data world, including watermarking, audit logs, anomaly detection, honeytokens, federated learning, confidential computing, and Trusted Research Environments. The governance section draws analogies to existing U.S. oversight regimes, including NIH rules on clinical trials, laboratory biosafety, patient genetic data, and dual-use research of concern, and identifies what worked and what did not in each. The authors state their work is a "vocabulary and set of hypotheses for further empirical validation" rather than a validated standard.

Why This Matters

The paper argues that in a world with widely accessible computational and coding resources, data controls may be among the most high-leverage interventions available to reduce the proliferation of concerning biological AI capabilities, because producing biological data remains costly and expertise-intensive while using AI is becoming cheaper. It also argues that controlling the data needed for adversarial fine-tuning could reduce concerns about releasing models openly, and that data standardization through BDLs could improve training data quality for legitimate applications. The paper notes that an international group of more than 100 researchers at the recent 50th anniversary Asilomar Conference endorsed data controls to prevent harmful AI applications such as bioweapons development.

Real-world applications:

  • Gene synthesis screening: the risk-scoring and provenance tools described map onto current practice for flagging sensitive gene synthesis order requests for human review.
  • Pandemic preparedness research: TREs would let vetted researchers work with high-risk pathogen data under supervision, supporting public health and pandemic preparedness work while limiting misuse.
  • Outbreak response: the authors recommend lifting BDL restrictions on pathogens causing a significant and ongoing disease outbreak, since real-time data sharing in those events can yield large public health benefits.
  • Health-data infrastructure transfer: the paper points to OpenSAFELY in England, the SAIL Databank's UK Secure Research Platform, the commercial platform DNAnexus, and the National COVID Cohort Collaborative (N3C) enclave in the United States as proof that the TRE model scales while preserving analytical utility, albeit outside large-scale AI training runs.

Industry relevance: AI developers training biological foundation models, cloud and hyperscaler providers that would host AI-scale secure compute, gene synthesis companies, diagnostics and pharmaceutical researchers, and data platform vendors all operate at points where this framework would apply. The authors acknowledge that running effective TREs for this domain is expensive and may require special compute clusters or partnerships with hyperscalers able to provide sufficient capacity.

Future Directions

  • Empirically validate the taxonomy. The authors call for systematic evaluation of how including or excluding particular datasets affects model capabilities, which would test whether functional data is as central as the framework assumes.
  • Convert tiers into objective, applicable rules. Specific viral families, data types, and data quantities for each BDL tier still need to be detailed, and the proposed Pathogen Data Board should revisit the framework regularly because AI and high-density pathogen data collection are advancing rapidly.
  • Fix unresolved technical problems. Watermark integrity through edits, synthesis, or lab work remains technically challenging; confidential computing to durably mask sensitive data during research is described as unsolved and possibly impossible; anomaly detection lacks good activity baselines; and decoys must be designed so legitimate researchers do not use them.
  • Build international alignment and shared infrastructure. The framework is described as workable long term only if all nations with significant biological research labs adopt similar rules, to avoid regulatory arbitrage, and Congress is urged to fund a compliant TRE platform whose ongoing operations could be supported through usage fees.

Target Audience

Policymakers and legislative staff working on AI or biosecurity; biosecurity and AI safety researchers; developers of biological foundation models and the data stewards who supply them; research funders and institutional oversight committees; gene synthesis and biodata platform providers; and legal or ethics scholars studying research oversight. The paper is also useful to scientists who produce pathogen data and need to understand how proposed classification rules would affect their work.

Authors’ abstract

Training data is an essential input into creating competent artificial intelligence (AI) models. AI models for biology are trained on large volumes of data, including data related to biological sequences, structures, images, and functions. The type of data used to train a model is intimately tied to the capabilities it ultimately possesses--including those of biosecurity concern. For this reason, an international group of more than 100 researchers at the recent 50th anniversary Asilomar Conference endorsed data controls to prevent the use of AI for harmful applications such as bioweapons development. To help design such controls, we introduce a five-tier Biosecurity Data Level (BDL) framework for categorizing pathogen data. Each level contains specific data types, based on their expected ability to contribute to capabilities of concern when used to train AI models. For each BDL tier, we propose technical restrictions appropriate to its level of risk. Finally, we outline a novel governance framework for newly created dual-use pathogen data. In a world with widely accessible computational and coding resources, data controls may be among the most high-leverage interventions available to reduce the proliferation of concerning biological AI capabilities.

Read the original paper