NVIDIA, Google DeepMind and a group of international research organizations have released a new open dataset containing predicted 3D protein-complex structures for more than 2,800 viruses. The structures are now available through the AlphaFold Database, giving researchers a large new resource for studying how viral proteins interact and potentially identifying targets for vaccines, diagnostics and treatments. NVIDIA is also releasing the GPU-accelerated BioNeMo Structure Prediction Pipeline used to generate the data.
The dataset covers 2,812 viral proteomes across 23 virus families considered relevant to human health. According to the AlphaFold Database, the project produced 5,279 high-confidence heterodimer structures, which contain two different interacting proteins, and 2,749 high-confidence homodimers formed from two copies of the same protein. The work includes viruses ranging from common-cold pathogens to emerging threats such as Mpox.
Researchers generated the structures with AlphaFold2 and AlphaFold-Multimer, using NVIDIA’s BioNeMo Inference Runtime to scale the prediction process across thousands of viral proteomes. Rather than examining proteins only as isolated molecules, the project focuses on complexes, where multiple proteins interact to perform biological functions. Those interactions can be especially important in infectious disease research because drugs, antibodies and vaccines often work by interfering with specific molecular structures or interactions.

Around 30% of the protein interactions being added have structures that were not previously documented in the Protein Data Bank, according to NVIDIA. That does not mean the AI-generated structures have been experimentally confirmed; the new entries are predictions and are labeled with confidence information so researchers can assess their reliability. High-confidence candidates can then be prioritized for laboratory experiments, potentially reducing the number of structures scientists need to investigate from scratch.
The release is part of a wider Pandemic Preparedness Portal now available through the AlphaFold Database. The portal combines newly generated viral complexes with other viral structure predictions already held by the database, giving researchers a single location for exploring proteins and interactions associated with viruses relevant to human health. AlphaFold DB itself now contains more than 260 million protein and protein-complex predictions.
NVIDIA is also making the BioNeMo Structure Prediction Pipeline openly available so researchers can generate structure predictions for their own protein targets. The GPU-accelerated workflow takes protein sequences through the steps needed to produce predicted 3D structures at scale. That is particularly useful for research groups studying pathogens that may not already be extensively represented in public structural databases.
The project brings together organizations including NVIDIA, Google DeepMind, EMBL’s European Bioinformatics Institute, the Coalition for Epidemic Preparedness Innovations, Seoul National University, Sungkyunkwan University, the Swiss Institute of Bioinformatics and the University of Glasgow. The collaboration combines AI infrastructure, biological databases and virology expertise rather than relying on a single company or laboratory.
The broader goal is to reduce the amount of basic structural work that would need to begin from zero during a future outbreak. During COVID-19, researchers benefited from decades of earlier coronavirus research that already provided substantial knowledge about viral proteins and their behavior. A newly emerging pathogen may not come with the same scientific head start, making precomputed structural information potentially valuable for generating early hypotheses.
The dataset does not replace experimental structural biology, and predicted structures still require validation before being used as definitive biological evidence. Its value is instead in narrowing the search space: researchers can examine likely protein interactions, compare related viruses and identify promising targets before committing to slower laboratory techniques. By making both the predictions and the generation pipeline openly accessible, the collaboration is aimed at giving researchers worldwide a larger starting point for future infectious-disease work.

