New specifications for submitting nucleotide sequence data

New INSDC standards support simpler submissions and improve synchronisation of global genome sequence data

The International Nucleotide Sequence Database Collaboration (INSDC) has released a new set of minimal specifications designed to modernise how nucleotide sequence data are described and exchanged globally.  

The INSDC is a global partnership between the European Nucleotide Archive (ENA) at EMBL’s European Bioinformatics Institute (EMBL-EBI), the National Center for Biotechnology Information (NCBI) and the DNA Data Bank of Japan (DDBJ). Together, these organisations have been collecting, storing, and openly sharing nucleotide sequence data since 1987. They also synchronise their records so that submissions to any one partner are available through all.

The updated specifications provide a unified framework that defines how sequence data and associated metadata should be submitted across the collaboration with the vision of accepting new partners in the future. By establishing these clear, baseline expectations, the INSDC lowers the barrier for future global collaborators to integrate their data seamlessly into the existing network.

What’s included in the INSDC’s new minimal specifications?

Over decades of advances in sequencing technologies, the scope and diversity of data submitted to INSDC have grown significantly. While this has enabled exciting new science, it has also made data submission a more complex process.

The new INSDC specifications set out clear expectations for how nucleotide sequence data and metadata must be structured in order to be accepted and shared across INSDC members.

These specifications define:

  • Which data types are supported, including sample types, raw sequence data, analyses, annotations and genome assemblies
  • What minimum information must be provided for each type. For example, required sample metadata or sequencing details
  • How these data connect together, linking biological samples to sequencing experiments and downstream assemblies
  • Agreed quality checks, ensuring that data submitted to one INSDC partner can be reliably understood and reused by all others

Overall, this means clearer submission requirements for researchers and a consistent structure for sequence records across the INSDC. 

What this means for ENA users

For ENA users, the new specifications provide a more transparent and consistent framework for submitting and reusing data. ENA will align its submission systems with the agreed INSDC standards, supporting faster and more consistent processing of newly submitted datasets. While most changes will happen behind the scenes, users can expect improved interoperability between ENA and its INSDC partners, making it easier to discover and reuse data across the databases.

The updated specifications also support ENA’s ongoing work to modernise services and prepare for future data types, providing a stable foundation for long-term data reuse.

Supporting global participation

The new specifications also support INSDC’s broader commitment to welcoming new members into the collaboration in order to better represent the global community of data providers and users. In welcoming new members, the INSDC aims to ensure that diverse perspectives shape the future direction of the collaboration. By clearly defining minimal expectations, the specifications help current and future partners understand what is required to contribute data within the collaboration.

The specifications will be maintained and updated over time in response to user feedback and emerging data needs, with new versions announced as they become available.

Funding

This work was supported by the Biotechnology and Biological Sciences Research Council (BBSRC) Metagenomics Portal IV [grant number BB/V01868X/1], in part by the Wellcome Trust [Darwin Tree of Life 226458/Z/22/Z]. Funding from the Gordon and Betty Moore Foundation through Aquatic Symbiosis [MOORE-8897]. The research reported in this publication was supported by the National Institute of Allergy and Infectious Diseases of the National Institutes of Health under award number U24AI183840 (Pathogen Data Network). The project was also supported by the Novo Nordisk Foundation and Wellcome via the AEGIS grant [NNF24SA0092560]. 

Funding was also provided for this project from the European Molecular Biology Laboratory. 

Edit

Related links

Tags: bioinformatics, cochrane, data, data standards, embl-ebi, ena, genomics,