The Pulse
NVIDIA and Partners Release Protein Predictions for 2,812 Viral Proteomes
NVIDIA, Google DeepMind and research partners have added predicted protein complexes from 2,812 viral proteomes to the AlphaFold Database. The release includes an open GPU-accelerated pipeline and predictions that researchers can test exper

AI.info Team ·
NVIDIA and a group of research organizations on Sept. 24 released predicted structures for viral protein complexes spanning 2,812 viral proteomes. The predictions are available through the AlphaFold Database’s Pandemic Preparedness Portal, alongside an open-source pipeline for generating structures from protein sequences. The release targets viruses relevant to human health, including families associated with common colds and Mpox.
More than 8,000 high-confidence complexes
The portal’s September update lists 5,279 high-confidence heterodimers and 2,749 high-confidence homodimers generated in the collaboration, covering 23 virus families relevant to human health. Heterodimers pair different proteins; homodimers consist of two copies of the same protein. The portal also includes 4,681 high-confidence viral homodimers from a separate dataset produced by the Atkinson Lab, so those predictions are distinct from the collaboration’s totals.
The effort brings together EMBL’s European Bioinformatics Institute, Google DeepMind, NVIDIA, the Steinegger, Mirdita and Grove labs, Philippe Le Mercier, and the ViralZone team at the Swiss Institute of Bioinformatics. The new predictions join an AlphaFold Database that says it now contains more than 260 million protein and protein-complex predictions.
Why researchers need complexes, not just single proteins
Many proteins function through interactions with other molecules. Those interactions can shape viral processes and may offer targets for medicines or vaccines, but structural information is missing for many viruses. The NVIDIA post says about 30% of the protein interactions in the new dataset have not been documented in the Protein Data Bank, which stores experimentally determined protein structures.
“What we’re trying to do is stockpile some of that knowledge ahead of time,” said Joe Grove, a professor of molecular virology at the Medical Research Council-University of Glasgow Centre for Virus Research and a project collaborator. The aim is to give scientists structural hypotheses to investigate before an outbreak, rather than waiting for experimental knowledge to accumulate after one begins.
Predictions are starting points for experiments
The collaboration used AlphaFold2 and AlphaFold-Multimer to predict structures, with NVIDIA’s BioNeMo Inference Runtime helping scale the computational work, according to the AlphaFold Database portal. NVIDIA also released its BioNeMo Structure Prediction Pipeline so researchers can run the workflow on their own protein targets.
These models are predictions, not experimentally confirmed structures. The portal labels them by confidence, and NVIDIA says researchers can test high-confidence predictions with laboratory methods. Chris Dallago, applied research science team lead in digital biology at NVIDIA, described the database as “an engine for hypothesis generation.”
An open resource for outbreak research
The dataset is part of the AlphaFold Database and is presented in a portal that gathers viral structure predictions in one place. The portal says the data are available for academic and commercial use under a CC BY 4.0 license, with attribution expected. NVIDIA’s announcement coincides with a United Nations General Assembly meeting on pandemic prevention, preparedness and response convened by the World Economic Forum in New York.
For researchers, the immediate deliverables are the downloadable predictions and the released pipeline. They can examine predicted interactions across viruses, choose candidates for follow-up, and compare the models with experimental findings; the structures themselves do not establish how a virus behaves in an infected organism.