Classification of pre-clinical data could boost drug research

Annotation of substantial dataset of 135 000 animal model and phenotype assays to help identify new drugs and drug targets

About the study

  • Data from in vivo assays have been grouped by disease or phenotype
  • The classification aims to make data more accessible to researchers
  • The findings could help scientists searching for new drugs

October 23, Hinxton – ChEMBL is a manually curated chemical database of bioactive molecules with drug-like properties, maintained by EMBL’s European Bioinformatics Institute (EMBL-EBI). ChEMBL is an important open access resource used for the discovery of new drugs. To further enhance the utility of the data in ChEMBL, researchers from EMBL-EBI undertook an intensive organisation task. They looked into a substantial dataset of more than 135 000 in vivo assays (experiments in whole animals) and grouped them by animal disease model or phenotypic endpoint.

“Our task was to annotate data from animal models and to group similar assays together. For example, we grouped all assays relevant for Parkinson’s disease, pain, inflammation, or diabetes, making it easier for a scientist interested in a certain disease to get a more comprehensive view of the data available,” explains Fiona Hunter, Biological Data Analyst at EMBL-EBI.

The results of this classification exercise, which has been published in the journal Scientific Data, could help researchers find patterns and uncover new evidence that can be useful for the discovery of new medicines. The annotation included three levels of classification and MeSH (Medical Subject Headings) terms – a controlled vocabulary thesaurus used for indexing articles in scientific publications – to make the format more useful and searchable.

“In the era of big data and artificial intelligence, the more data available the better, but we must be careful not to combine data that is incompatible,” adds Andrew Leach, EMBL-EBI Head of Chemistry Services. “We hope that this initiative will make it easier for users to know which data can be reasonably well combined and integrated”.

Leach explained that the annotated dataset also plays a role in addressing potential safety and toxicological aspects of drug research, by helping to flag up potential issues, which could then be addressed in an earlier stage of the drug development. “Our broader goal is to help us better understand and interpret what might be ultimately observed in human patients in clinical trials”.

The annotated dataset has initially been made available for download via FTP but will subsequently be accessible as part of a later release of the ChEMBL database, and via the EMBL-EBI web interface or web services (https://www.ebi.ac.uk/chembl/ws).

Source article

Hunter FMI, et al. (2018) A Large-Scale Dataset of in vivo Pharmacology Assay Results. Sci. Data. 5:180230 DOI: 10.1038/sdata.2018.230.

Funding

This work was supported by the European Union’s Seventh Framework Programme (FP7) under grant agreement no 602156, the Strategic Awards from the Wellcome Trust (104104/A/14/Z) and by core funding from the European Molecular Biology Laboratory (EMBL).

 

Edit

Tags: animal models, annotation, assays, biological data, ChEMBL, ChEMBLdb, drug discovery, drugs, phenotype,