InterPro 105.0: AI for protein classification

InterPro 105.0 is now live. This AI-driven update makes it easier than ever to explore the protein universe
InterPro logo
InterPro: new release. Image credit: Karen Arnott/EMBL-EBI

The InterPro 105.0 release marks one of the most substantial updates to the InterPro data resource in years, bringing together cutting-edge AI annotations, structure-based domain insights and expanded viral proteome coverage in one unified resource.

First, InterPro-N delivers 1.8 billion deep-learning-driven protein annotations, boosting UniProtKB coverage to over 90% and filling long-standing gaps in domain classification. Next, TED domains harness AlphaFold-derived models to define independent folding units, allowing direct comparison of structure-based and sequence-based annotations. Finally, through BFVD integration, researchers can explore over 300,000 high-confidence viral protein structures alongside traditional InterPro annotations.

InterPro-N: AI-driven protein annotations

Since 2021, the InterPro team have been collaborating with Dr. Lucy Colwell’s team at Google DeepMind to extend protein domain classification with deep learning. Initially launched as Pfam-N, a collection of Pfam matches predicted using deep learning, the approach has evolved into InterPro-N (read the post on the Pfam blog), developed by Dr. Laetitia Meng-Papaxanthos and Dr. Felipe Llinares-López. Inspired by segmentation techniques used in computer vision, notably the Maskformer architecture, this model was trained using data from all 13 InterPro member databases. Initially, only Pfam annotations predicted by this model were available within the Pfam-N section of the domain viewer on the InterPro website.

Today, we are excited to announce that all 1.87 billion InterPro-N predictions are now fully available on the InterPro website, via the REST API, and downloadable from the EMBL-EBI FTP site.

InterPro-N has been trained on InterPro 105.0 data, with inferences made against UniProtKB 2025_02, ensuring that both InterPro and InterPro-N are in perfect sync.

The remarkable increase in the number of sequences in UniProtKB over recent years has meant that although 84% of sequences are annotated with traditional InterPro member database signatures, a large fraction of sequences still lack definitive classification into domains and families. InterPro-N addresses this gap by significantly increasing the number of sequences that receive functional annotations. Indeed, over 90% of UniProtKB sequences now receive at least one annotation, either from traditional InterPro methods or through the expanded reach of InterPro-N.

Evolution of the sequence coverage of InterPro over the past 5 years. In InterPro 105.0, 84% of UniProtKB sequences are hit by at least one member database signature, and nearly 90% of sequences are annotated with an InterPro-N prediction, significantly increasing the coverage of the known protein universe.
Number of UniProtKB sequences matched by InterPro and InterPro-N for each member database. InterPro-N consistently increases the coverage of UniProtKB across all member databases in InterPro.

By default, the website displays traditional InterPro matches, supplemented by novel InterPro‐N matches. When both methods report a match, the InterPro annotation is retained unless the InterPro‐N one is at least 5% longer. It is possible to change this behavior via the options menu, accessible by clicking the Options button located above the protein domain viewer. Three alternative display modes are available:

  • InterPro: display only the traditional InterPro annotations
  • InterPro-N: display only InterPro-N predictions
  • Stacked: when both InterPro and InterPro-N report a match, both annotations are displayed vertically stacked. This is the optimal display mode for visual comparison of coexisting InterPro and InterPro‐N matches, but please note it may result in a crowded view of the domain viewer showing a high number of annotations.

The selected display mode is remembered by the web browser and will persist across all pages displaying matches for a sequence that has been annotated using InterPro-N. However, please note that InterPro-N predictions are not yet available for PDBe chain sequences, protein isoforms, or sequences submitted through the InterProScan web search.

Screenshot of InterPro and InterPro-N annotations for Transcription factor otaR1 (A0A1R3RGK4) using the ‘Stacked’ display mode. InterPro-N annotations are distinguished by a leading sparkles icon (✨). Within the Families section, only InterPro reports a PANTHER family. In the Domains section, InterPro and InterPro-N largely concur, showing only slight variation in boundary assignments. Notably, while the bZIP domain in InterPro (IPR004827) integrates five signatures (two from PROSITE, two from Pfam, one from SMART), only the two PROSITE signatures are detected by InterPro for this protein. In contrast, Interpro-N identifies two additional signatures, demonstrating its enhanced sensitivity.

TED domains

This release also features the integration of The Encyclopedia of Domains (TED), developed at University College London by the groups of Professor David Jones and Professor Christine Orengo. TED provides a large-scale classification of structural domains derived directly from AlphaFold protein structure predictions. It identifies potential independent folding units using multiple domain boundary prediction algorithms. Within InterPro, TED domain predictions are displayed alongside matches from member database signatures (e.g. Pfam, CDD), allowing users to compare sequence-based annotations with structure-based domain insights.

Screenshot of TED domains, InterPro matches, and InterPro-N predictions for Pestheic acid cluster transcriptional regulator 1 (A0A067XNJ0). InterPro reports CATH-Gene3D and SUPERFAMILY domains that agree with the TED domain at the N-terminal, while the CATH-Gene3D and SUPERFAMILY annotations predicted by InterPro-N agree on the C-terminal TED domain, in particular the discontinuous CATH-Gene3D domain.

Integrating predicted viral proteins structures from BFVD

The AlphaFold Protein Structure Database (AFDB) contains a vast repository of predicted structures, yet fewer than 4,000 of nearly six million viral proteins in UniProtKB have predicted structures. To address this, the group of Martin Steinegger developed the Big Fantastic Virus Database (BFVD), which houses over 300,000 protein structures predicted using ColabFold applied to the viral sequence representatives of the UniRef30 clusters. In InterPro 105.0, 8,136 entries include at least one BFVD prediction.

Users can access BFVD predictions either directly on the protein page or through dedicated links from InterPro entries and member database signatures.

Screenshot of the BFVD predicted structure of Replicate polyprotein P2AB (P21405).

Other website updates

Extending representative domains

With InterPro 96.0 (Sep 2023), we introduced a track for representative domains. These domains are automatically selected to maximise sequence coverage while minimising overlap, thereby improving the clarity and conciseness of protein domain visualisations. This approach provides a clearer overview of domain organisation within a sequence.

To be eligible for selection as a representative domain, a match must originate from a member database signature classified as either a domain or a repeat. Initially, only signatures from the following member databases were considered: Pfam, CDD, PROSITE Profiles, SMART, and NCBIFAM. With this release, domains from CATH-Gene3D and SUPERFAMILY are also now eligible.

Screenshot of human Cellular tumor antigen p53 (P04637). (A) The representative domains track before InterPro 105.0, showing only the Pfam P53 DNA-binding domain, while other domains are known to exist. (B) The representative domains track and all other domain tracks in InterPro 105.0. The central DNA-binding core domain is now represented by a CATH-Gene3D match because it is slightly longer. Additionally, CATH-Gene3D matches are used for the N-terminal and C-terminal regions. In the N-terminal area, two Pfam matches could more accurately represent the two activation domains (AD1 and AD2); however, since the underlying Pfam signatures are classified as conserved sites, they are ineligible for selection as representative domains. We plan to address this in future InterPro releases.

Grouping PDBe structures and chains by sequence

Protein structures from PDBe often consist of multiple chains, each representing an individual molecular entity, such as a protein subunit. Although each chain is defined by its sequence, many structures contain several chains that are identical copies of the same protein. Displaying an InterPro match viewer for every chain can therefore lead to redundancy. To address this issue, we now display one viewer per unique sequence, ensuring that each distinct set of functional and structural annotations is shown only once, even when it appears in multiple chains. For example, the structure of a Saccharomyces cerevisiae Ipl1 peptide Bound to dwarf Ndc80 complex (8v11) has eight chains but four distinct sequences; therefore, only four protein viewers are displayed.

Improving the integration of DisProt and RepeatsDB

DisProt is a comprehensive database that curates experimental evidence of protein disorder from scientific literature. DisProt annotations, when available, are shown in the domain viewer on the InterPro website. However, before this release, DisProt annotations were in a separate section (External Sources) at the bottom of the viewer, and each region was displayed as a separate track. In this release, DisProt annotations are displayed in the Intrinsically Disordered Regions section with disordered regions predicted by MobiDB-lite. Additionally, the DisProt regions are colored based on their structural aspect, as defined by the Intrinsically Disordered Proteins Ontology (brown: structural state; purple: structural transition; red: disorder function).

Screenshots of the intrinsically disordered regions, provided by MobiDB-lite and DisProt, in 7SK snRNA methylphosphate capping enzyme (Q7L2J0). (A) Before this release. (B) From this release onwards.

RepeatsDB is a specialised resource that annotates tandem repeat structures in proteins, which are often linked to specific functional or structural properties. Previously, RepeatsDB annotations were also placed in the External Sources section, making them less accessible for visual inspection. Moreover, the annotations did not display individual units; instead, the repeated region was shown as a single contiguous block. In this release, RepeatsDB annotations have been relocated to the Domains section, allowing for more intuitive visual integration with other protein domain annotations. Additionally, individual repeated units are now displayed in red and blue, with any insertions marked in yellow.

Screenshots of the RepeatsDB annotations in Pyruvate carboxylase (A0A0H3JRU9). (A) Before this release. (B) From this release onwards.

Data updates

This update includes 342 new entries, increasing the total number of InterPro entries to 48,003. In addition, since our last release, 385 member database signatures have been integrated.

As with every new version, the underlying UniProtKB sequences have been refreshed. InterPro 105.0 is based on UniProtKB 2025_02. While the overall sequence coverage has remained relatively stable, our continuous efforts ensure that the quality and relevance of the annotations are always improving. The following table summarises recent InterPro releases:

InterPro versionRelease dateUniProtKB versionNumber of sequencesCoverage (%)
102.003 Oct 20242024_05248,838,88783.95
103.028 Nov 20242024_06254,254,98783.98
104.006 Feb 20252025_01253,206,17184.02
105.024 Apr 20252025_02252,761,75284.04

In this release the following member databases have been updated:

  • HAMAP 2025_01 (2 new signatures)
  • Pfam 37.3 (348 new signatures)
  • PROSITE patterns 2025_01 (no new signatures)
  • PROSITE profiles 2025_01 (20 new signatures)

The InterPro coverage of a few key species is represented on the graph below.

Sequence coverage of key species. Sequences matched by at least one signature integrated in an InterPro entry are represented in green, sequences matched by signatures not yet integrated are represented in blue, and sequences not matched by any signatures are represented in grey.
Edit

Related links

Tags: artificial intelligence, bioinformatics, database, embl-ebi, interpro, machine learning, protein,