Happy anniversary, Swiss-Prot!

A unique collaboration that unifies the world of protein data: reflections on UniProtKB/Swiss-Prot, 30 years on.
Swiss-Prot 30th anniversary
Celebrating 30 years of Swiss-Prot

The Universal Protein Resource Consortium provides the global research community with a gold-standard knowledgebase: UniProtKB/Swiss-Prot, a trusted source of freely available protein information that has played a pivotal role in the rise of digital biology. For the past 30 years, the people working on the resource have strived to provide the highest-quality data.

Connecting biology

With close to 800,000 visits every month, UniProt has become a foundation of research in biology and molecular medicine. For many scientists it is often the first step in their discovery journey – a virtual crossroad where their direction of research may be decided.

“Over the years, many people have come to rely heavily on our resources,” says Alex Bateman, Head of Protein Sequence Resources at EMBL-EBI and the Principal Investigator for UniProt. “Everyone gets unrestricted access to this data and a fair chance to use it to advance research, regardless of where they come from.”

Many users shape their research based on information they find in UniProt, and build a case drawing on its extensive analysis services and high-quality annotations. This practice has the added benefit of cutting research time in half, as reported in a large-scale user survey carried out in 2014. The resource has been mentioned in over 15 000 patent citations since 2000.

Keeping pace with demand

To support its seamless integration into research practice, UniProt relies on a complex infrastructure managed by its development team, a large group of software engineers and bioinformaticians based in Switzerland, the UK and the US. This technical infrastructure supports the vital biocuration activities that make it such a reliable source of information. UniProtKB/Swiss-Prot biocurators personally read and select the most relevant papers, annotating protein data records to ensure they are presented in context and offer structured information about their influence in a living system.

“All of the variations between us, in how we are, how we look and behave, our health, have a lot to do with variations on the protein level,” explains Claire O’Donovan, Protein Function Content Team Leader at EMBL-EBI. “Being able to track down this data and connect it with consequence in the era of Big Data is what we are after. Ultimately, we would like UniProt to be sophisticated enough to predict relationships between genes and proteins without the core science having to be done first in the lab.”

UniProt is growing at a tremendous pace: launched in 2002, close to 80% of today’s sequence content has been collected over the past three years alone. The UniProt Consortium works with authors and publishers to improve the annotation process, pushing to raise industry standards.

“We receive a huge amount of direct input from authors, and from people submitting feedback to help us improve our services," says Alex. "To say we appreciate this input would be a huge understatement.”

The origins of UniProtKB/Swiss-Prot

UniProt is an example of thriving international teamwork across three sites with over 100 employees. Its beginning go back to 1986, when Amos Bairoch started developing programs for protein research at the University of Geneva. While distributing his sequence analysis software with PIR-PSD, the world’s first collection of protein sequences chronicled in a series of atlases from 1965 to 1978 by Margaret Dayhoff, Amos ran into difficulties with cross-referencing, annotations and vital characterisation information due to the PIR-PSD format.

Inspired by EMBL’s Nucleotide Sequence Data Library, Amos released ‘PIR+’ in a similar format, later renaming it Swiss-Prot. In 1986 he distributed the first release of the database freely: Swiss-Prot version one contained 3,900 sequences. Today, the UniProt Knowledgebase offers over 63 million.

EMBL’s data library and Swiss-Prot were a natural fit, and led to a lasting collaboration. EMBL agreed to distribute the database and participate in Swiss-Prot’s maintenance, and in 1987 Dr Rolf Apweiler, now Director of EMBL-EBI, was hired as the first EMBL Swiss-Prot curator. In 1994, he took on the leadership of the growing EMBL Swiss-Prot group. 

What’s next?

“Clinical research could benefit a lot from the progress we’ve made in research, establishing data standards and agreeing on a shared language,” says Maria Martin, team leader for Protein Function Development at EMBL-EBI. “Without these tools, it is much more difficult to put together new and existing data in useful ways, for example to explain or predict the efficiency of a drug.”

The UniProt Consortium works with a broad range of life-science and biomedical research communities to ensure they are ready for new technologies, data types and even working styles. The long-standing collaboration between Switzerland, the UK and the US gives the resource the stability it needs to keep high-quality protein data freely available, and to proactively develop to meet future needs.

Find out more

Read more about how UniProt works and where it is headed

Read about how Swiss-Prot got started, and the origins of UniProt

Connect with EMBL-EBI at Biocuration 2016

Edit

Tags: Alex Bateman, Claire O'Donovan, Maria Martin, Swiss-Prot, UniProt, UniProtKB,