First bulk CRAM submission to ENA
The first large-scale read data set in the CRAM compressed format has been submitted, undergone processing and been made public at ENA. Using CRAM in lossless mode, the submission represents a pre-publication data release from the Wellcome Trust Sanger Insitute and comprises around 4,000 run records covering a number of pathogen species. Data are available for download in both CRAM and FASTQ formats. The data set in CRAM format consumes 80% of the disk space or network bandwidth for download required for its gzipped FASTQ equivalent.
An example of a study in this data set can seen here.
Further information on CRAM sequence data compression technology is available here. The starting point for data submissions to ENA, including those in CRAM format, is here.
Edit