Using deep learning to annotate the protein universe
Nature Biotechnology 21 February 2022
0.1038/s41587-021-01179-w
MGnify are thrilled to announce the long-awaited new release of their protein sequence database. At 2.4 billion non-redundant sequences, this is more than double the size of the 1.1 billion sequence previous release. Part of the delay in getting this release out is due to a complete reengineering of our processes to generate and store this data. Specifically, we have moved to a relational database model, which allows us to improve the linking of the protein sequences to the assemblies and actual contig locations they are derived from.
As well as allowing us to provide information on which biome(s) a particular sequence is found in, crucially, this means that rich information on the genomic context of the protein sequence can now be provided. For example, you can find out if another gene sits upstream or downstream of the one coding for your protein of interest.
As 2.4 billion sequences can be tricky to handle, we have clustered the data (at 90% coverage and identity) into a much more manageable set of ~620 million clusters. All the information about the full list of sequences that are represented by each cluster is provided in our download files. Each cluster has a representative sequence, and Pfam annotations are provided by HMMER for these cluster reps.
We are particularly excited to make available a new set of annotations in this release provided by our colleagues at Google AI using ProtENN2. Building on the methods developed by Max Bileschi and colleagues, convolutional neural networks are used to annotate each residue (for all sequences in the MGnify database) with a Pfam family or clan label, which are then converted into domain calls. We estimate that ProtENN2 has predictions for an additional ~200 million proteins that Pfam/HMMER does not have any annotations for at Pfam gathering thresholds.
A paper on this method is forthcoming; please direct inquiries to mlbileschi@google.com.
As the vast majority of proteins identified from metagenomics datasets and included in this release are not represented in the reference databases, these new approaches to understanding and annotating them are vital in our efforts to shed light on this vast field of protein dark matter.
You can find all the data comprising this latest MGnify protein database release on the MGnify FTP site. We provide the fasta sequences (for the full set of 2.4 billion proteins as well as for the cluster reps), the ProtENN2 annotations, the HMMER annotations, the protein to assembly mappings, the biomes of the sequences, and a selection of other useful mappings and statistics. We would welcome your feedback on the release and the data we make available at any time.
EditNature Biotechnology 21 February 2022
0.1038/s41587-021-01179-w