ENA launches Data Hubs Portal

The Data Hubs Portal is a new interface that enables users to setup and manage pre-release and/or public Data Hubs at the ENA.
Decorative image showing ENA logo.
ENA logo

The European Nucleotide Archive (ENA) provides a comprehensive record of the world’s open nucleotide sequencing information. It resides amongst an ecosystem of databases and resources that all interconnect, and ultimately build on the ENA services and content. To share sequencing data to the ENA, users are pointed to Webin, with more personalised accounts, that don’t enable for close collaboration with colleagues.

The Data Hubs Portal is a new interface that enables users to setup, access, manage and search pre-release and/or public Data Hubs at the ENA. The Data Hubs have a number of unique features for sequencing data at the ENA:

  1. Collaboration: Data Hubs enable data sharing amongst a group of collaborators in pre-release, private mode, prior to publishing and data release.
  2. Configurations: Data Hubs enable researchers to use preferred tools for submission, search and retrieval of pre-release and/or public data and metadata.
  3. Role definitions: Data Hubs include (1) a coordinator – who sets up and manages the Data Hub; (2) provider(s) – who share sequence data to the Data Hub; and (3) consumers – who only need access and can download data and metadata.
  4. Sustainability: Data Hubs are built on top of ENA infrastructure, reusing the same storage, data and metadata models.
Diagram showing how ENA Data Hubs can be used and what functionalities they provide.

Originally developed for foodborne infectious disease data under the COMPARE project, the Data Hubs were expanded to support COVID-19, and even further to priority infectious diseases under the Pathogen Data Hubs. Now, the Data Hubs are available for the breadth of data across the ENA.

What’s new?

  • Accessibility: An interface for users to setup, access, manage and search Data Hubs.
  • Breadth: Going beyond just infectious disease data, to support any domain of data in the ENA.
  • Pre-release collaborative data sharing: ENA data submissions are currently dedicated to a single primary contact. With the Data Hubs Portal, a user can define individuals/groups sharing data as opposed to retrieving data, prior to data release.

An example

Just one example that helps to visualise when Data Hubs are useful is through the consortium concept. Different institutes within a consortium may be carrying out various activities. Collaborators A and B in two different institutes may be carrying out genomic sequencing and uploading data to repositories as part of data management. However Collaborators C, D and E based in other teams/institutes may be interested in analysing and gaining insights into the data.

Here a Data Hub enables the various collaborators to connect on pre-release data. A and B can share/provide data, where C, D and E can download/consume that data, analyse and in some cases even return the analyses back to the Data Hub. Once all parties are happy with the data and any associated publication has been released, then the Data Hub can be released.

What’s next?

The Data Hubs Portal will include further features in the future to enable user collaboration with Data Hubs, and improved documentation. Users have helped provide valuable feedback in the testing phase of the portal, and we welcome further feedback for users to help drive additional features for the future.

Acknowledgements

Research reported in this announcement was supported by the NIAID BRC program of the National Institutes of Health under award number U24AI183840, Pathogen Data Network.

The Data Hubs Portal project has received funding from the European Union’s Horizon Europe and Horizon 2020 research and innovation programmes under the following grant agreements:

Edit

Related links

Tags: bioinformatics, covid-19, embl-ebi, ena, FAIR data, genomics, infectious disease, open data, pathogen,