📣 Try out the NEW Gateway to Research and let us know what you think.

We're looking for users to test the new service during August and September and share their feedback. Express your interest by completing this short form.

Accelerating medicine development timelines through new approaches in knowledge extraction from diverse biological data sets

Lead Research Organisation: UNIVERSITY OF EDINBURGH
Department Name: Edinburgh Cancer Research Centre

Abstract

Never has the impact of the development time of a new medicine been better understood with the current situation in the UK relating to COVID. This remains and has always been a big driver for Pharmaceutical companies, to reduce the time it takes from identification of an interesting compound or vaccine to a marketed medicine. The main focus of this project is to develop new Artificial Intelligence and Machine Learning (AI/ML) analytics tools to speed up the identification of immunological medicines, including small molecules, biologics and vaccines.

One of the biggest challenges we face in drug development that delays our ability to bring medicines to market quickly is the high rates of candidate molecule termination due to poor clinical efficacy and poor pre-clinical safety. To reduce the high rates of compound attrition, we are implementing disease relevant and physiological human cellular models in early stage discovery. We are using these cellular models to identify new therapeutic targets and screen our pharmacological agents to prioritise those with a better chance of success and to stop projects earlier with either efficacy or safety liabilities. These assays often rely on cellular imaging, in recent years automated microscopy has opened up the ability to characterise the state and phenotypes of cells, at single cell and even subcellular resolution. Thus, biological diversity can be visualized and the effects of perturbations on cells can be quantified more richly than by almost any other means. Relating this back to the challenge of compound failure we can use this imaging data to study compound effects in complex human in vitro systems to enable us to identify potential efficacy and safety risks much earlier in drug discovery enabling us to prioritise those medicines with a higher chance of success. As well as automated imaging, advancements have also been made to study other high content data sets such as transcriptomic, and proteomic information. This used to be completed in a screening cascade of assays all run independently to optimise our pharmacology, we are now able to multiplex these endpoints in the same system enabling us to pull together a complete fingerprint of cellular activity much faster.
With these models in place, the current challenge in drug discovery is that the level of complexity of both cellular models as well as the data derived from these models is often beyond the level of data analytics technologies we have available. There is an interesting relationship between data volume and understanding. Initially if you increase the amount of information you have on a system, your understanding increases before reaching a point where understanding rapidly decreases and increased data actually causes confusion as it becomes impossible to interpret.
This project is focused on resolving this bottleneck in two phases. Firstly, in recent years, AI methods, specifically convolutional deep learning methods, have revolutionised the field of computer vision by achieving performance often better than humans in image interpretation. In phase one of this project we intend to use AI methods to build a toolbox of image analytics solutions that will be used across our cellular imaging studies, to gather more information from our high-resolution images and transform millions of pixels into parameters. Similar deep learning networks have also now demonstrated the ability to explore vast data sets to interpret and deliver biological understanding. Thus, the second phase of this project will be applying AI tools to transform the millions of parameters obtained from all our high content technologies into features and mechanistic information that can be used to support project decisions.
By utilising the immense power of AI approaches in image analysis and big data analytics we hope to enable a better understanding of both chemical and genetic perturbations, ultimately improving our success rate in the clinic.

Technical Summary

Early stage drug discovery involves conducting panels of well-validated assays to generate datasets on collections of candidate chemical, biological or genetic entities. Such datasets include metabolomic, proteomic, transcriptomic and image-based cellular profiles and provide decision-making data on the suitability of a candidate to be moved to clinical phases of development. Historically, these types of datasets have been used in isolation but growing evidence shows that by considering the same datasets in relation to one another through improved machine learning strategies, rates of attrition in drug discovery can be reduced; an effect mediated through identification of disease-associated phenotypes and mechanisms in early screening and greater fidelity in early prediction of efficacy and toxicity. To this aim the secondee will utilise immunological datasets derived from primary T cell and iPSC-derived macrophage assays to develop ontological frameworks within the GSK R&D information platform in order to integrate imaging and omics datasets and simultaneously provide a structure suitable for future reuse. The framework will provide inter-relation of previously disparate databases enabling the secondee to deploy deep learning and machine learning tools to map responses to fingerprints from compounds or disease profiles annotated by in vivo or clinical data. Critically, as the effect of drug candidates are highly dose-dependent, the analytics developed will trace dose across multidimensional space to inform upon pathway preferentiality, off-target effects, and action of multiple targets in each therapeutic mechanism. This will provide new ranking metrics for medicinal chemists to optimise a drug series and inform on opportunities for combination screening. The work will provide the secondee with valuable networking opportunities and insights across the pharmaceutical industry through collaboration with units including AI/machine learning and medicinal chemistry.

Publications

10 25 50
 
Description We have developed a fully opensource software pipeline that enables the integration of complex datasets pertaining to drug testing in non animal laboratory models. The software then enables exploration of the integrated data using new analytical methods such as artificial intelligence tools. Importantly all analysis is auditable ensuring a traceable link between the AI generated results, the methods used and the source data. The traceability is important for drug discovery data that may be used in regulatory submissions and intellectual property protection.

This project was part of an MRC innovation scholarship award to enable cross sector working between academia and industry and fill the skills gap in advanced computational approaches in drug discovery. Following the award the innovation scholar secured a full time industry position with and AI company operating in the life sciences sector.
Exploitation Route The Phenonaut software that we have published and provides underlying source code can currently be downloaded and implemented by academic and industry groups. The software could be further developed to increase it applicability and useability.
Sectors Healthcare

Pharmaceuticals and Medical Biotechnology

URL https://github.com/CarragherLab/phenonaut
 
Description The Phenonaut software developed as part of this program has been implemented within the pharmaceutical industry to advance the application of computational methods including Artificial Intelligence methods for drug discovery, GSK who was an industry partner on the award has been the first to implement this software and an example is published as a pre-print: https://arxiv.org/abs/2404.18960
First Year Of Impact 2024
Sector Pharmaceuticals and Medical Biotechnology
Impact Types Economic

 
Title Phenonaut 
Description This is a novel software platform for analysis of multiomic and single-omics datasets 
Type Of Material Improvements to research infrastructure 
Year Produced 2023 
Provided To Others? Yes  
Impact Phenonaut fills an unmet need in multiomics enabled phenotypic drug discovery, allowing integration of multiomic data and application of machine learning, and data science techniques for phenotypic space exploration, hit calling and prediction. We anticipate Phenonaut will be utilized by a broad user base to enable complex multiomic data analysis, which is robust and audited to support both basic and translational research applications, thus filling a significant gap in currently available tools. The software is currently being deployed in GSK and the Carragher lab with further functionality and publications planned. Source code is publicly available in github 
URL https://github.com/CarragherLab/phenonaut
 
Description GSK staff secondment 
Organisation GlaxoSmithKline (GSK)
Department Research and Development GSK
Country United Kingdom 
Sector Private 
PI Contribution Supervision, datasets and intellectual guidance
Collaborator Contribution Senior staff supervision, computational resources and datsets.
Impact Talk was presented by Dr Steven Shave at the ELRIG Drug Discovery 2022 meeting 4th/5th October Excel London. A manuscript titled: "Phenonaut; multiomics data integration for phenotypic space exploration" is currently under review in the journal oxford bioinformatics. Source code is publicly available on github: https://github.com/CarragherLab/phenonaut This collaboration is Multidisciplinary: Cell biology; drug discovery; computational biology; Artificial Intelligence/Machine Learning; software development.
Start Year 2021
 
Description CRUK Scottish Cancer conference 
Form Of Engagement Activity A talk or presentation
Part Of Official Scheme? No
Geographic Reach National
Primary Audience Policymakers/politicians
Results and Impact Presented at breakout workshop session titled: "University of Edinburgh: A nationwide collaboration addressing Scotland's cancer health challenge" at the Scottish Cancer Conference held at the Edinburgh International Conference centre on 24th November 2024. The purpose was to introduce the CRUK Scotland Center, highlight thematic research area priorities including Brain tumour and bench-to-bedside translational research case studies followed by panel discussion.
Year(s) Of Engagement Activity 2024
URL https://www.scottishcancerconference.org.uk
 
Description Laboratory visit from Scott Arthur, MP to support the rare cancers bill 
Form Of Engagement Activity Participation in an open day or visit at my research institution
Part Of Official Scheme? No
Geographic Reach Local
Primary Audience Policymakers/politicians
Results and Impact I hosted a visit and tour of our drug screening laboratories at the Institute of Genetics and Cancer from Scott Arthur, MP for Edinburgh South West (10th April 2025). We discussed the latest chalenges and opportunities in brain cancer research. This was to support the Private Members' Bill "Rare Cancers Act 2026" being led through UK parliament by Scott Arthur. The aim of the Bill is to make provision to incentivise research and investment into the treatment of rare types of cancer (including brain cancers).
Year(s) Of Engagement Activity 2025
 
Description Presentation at annual AACR meeting. 
Form Of Engagement Activity A talk or presentation
Part Of Official Scheme? No
Geographic Reach International
Primary Audience Professional Practitioners
Results and Impact AACR ANNUAL MEETING 2023, Orlando, Fl, April 14th-19th, 2023. Talk Title: "High content phenotypic and pathway profiling identifies novel drug mechanisms-of-action tailored towards cancers of unmet need".

My talk was presented at the education day session to share latest advances in cell based screening technologies which promise to expedite drug discovery at reduced cost and increase clinical translation success rates in complex cancers of unmet need.
Year(s) Of Engagement Activity 2020,2023
URL https://www.aacr.org/wp-content/uploads/2023/03/AACR2023_ProgramGuide_032823.pdf
 
Description Presentation to CRUK business beats cancer Edinburgh annual event 
Form Of Engagement Activity A talk or presentation
Part Of Official Scheme? No
Geographic Reach Regional
Primary Audience Industry/Business
Results and Impact Represented Cancer Research UK Scotland centre by presenting our drug discovery work at the "Business Beats Cancer Edinburgh Gala Dinner" event held on 11th May 2023, Prestonfield Hotel, Edinburgh (300+ specially invited prominent business leaders and contacts attending). The primary purpose was to raise awareness of our advanced drug discovery capabilities and raise funding for Cancer Research UK. The event raised £120k for cancer research and a follow up lab visit from board members and supporters took place of 7th February 2024
Year(s) Of Engagement Activity 2023
URL https://www.cancerresearchuk.org/business-beats-cancer-edinburgh
 
Description Public-Patient focussed Brain Cancer Research showcase 
Form Of Engagement Activity A talk or presentation
Part Of Official Scheme? No
Geographic Reach Local
Primary Audience Patients, carers and/or patient groups
Results and Impact I presented our research work to the public (>80 attendees) at the EDINBURGH BRAIN CANCER RESEARCH SHOWCASE organized by local brain cancer oncologist DR Sara Erridge and held at the Institute of Genetics and Cancer (27th March 2025).

The purpose was knowledge exchange with public and patient groups
Year(s) Of Engagement Activity 2025