Longitudinal machine learning of molecular and phenotypic trajectories of pulmonary hypertension.
Lead Research Organisation:
IMPERIAL COLLEGE LONDON
Department Name: National Heart and Lung Institute
Abstract
Our vision is to fundamentally redefine the diagnosis and treatment of Pulmonary Hypertension (PH), a multifaceted disease that carries significant diagnostic and prognostic difficulties. Even though there are numerous cross-sectional studies on PH, they cannot be combined because the patients are all at different stages of the disease along with different comorbidities. By integrating high-dimensional molecular profiling into traditional longitudinal cohort studies, we aim to leverage machine learning, genomics, and vascular biology to predict future molecular and clinical measures at any time point in a patient's journey. This approach can potentially discover biological mechanisms that drive disease progression and identify biomarkers in patients even when they cannot visit the clinic to provide data.
Aims and Outcomes:
1. Develop machine learning models to predict whole transcriptomes from blood biopsies at any point in a PH patient's journey. This could enable a new form of molecular classification for PH at previously unexamined time points, offering more precision than current clinical classifications.
2. Create an efficient computational system for real-time analysis of high-dimensional data, such as thousands of genes, thus overcoming the current methodological gap for genomic data captured at multiple times.
3. Extrapolate molecular changes observed in blood biopsies to changes in the pulmonary vasculature, providing a non-invasive method for investigating disease mechanisms.
4. Determine the optimal sample size and design for longitudinal molecular studies, increasing these studies' cost-efficiency and statistical power for other diseases.
Our interdisciplinary approach begins with computational modelling of gene expression and symptom trajectories across multiple years following a patient's diagnosis. We will begin by combining molecular trajectories with baseline factors, including gender, ethnicity, and socio-economic status. Time-dependent changes in blood transcriptome and methylome will be associated with changes in patient symptoms, such as mean arterial pressure and 6 minute walk distance. Historically measured gene expression profiles will be used directly as longitudinal inputs in our machine learning models to predict future molecular profiles associated with pulmonary vascular remodelling and patient outcome.
To ensure scalability to many patient measures and time points, we will evaluate and extend a machine learning technique called Gaussian Processes, which has successfully modelled lower-dimensional longitudinal data such as body weights and CO2 emissions. The algorithm will be trained through an iterative process, by gradually increasing the amount of longitudinal data from patients' molecular profiles and electronic health records. Our multidisciplinary team of computer scientists, epidemiologists, biologists and clinicians, will enable the algorithm to use biologically relevant features and be scalable for clinical applications.
The data and methods from this project will not only provide PH patients with better information about their disease progression, but can also be used to understand other complex diseases and age-related disorders.
Aims and Outcomes:
1. Develop machine learning models to predict whole transcriptomes from blood biopsies at any point in a PH patient's journey. This could enable a new form of molecular classification for PH at previously unexamined time points, offering more precision than current clinical classifications.
2. Create an efficient computational system for real-time analysis of high-dimensional data, such as thousands of genes, thus overcoming the current methodological gap for genomic data captured at multiple times.
3. Extrapolate molecular changes observed in blood biopsies to changes in the pulmonary vasculature, providing a non-invasive method for investigating disease mechanisms.
4. Determine the optimal sample size and design for longitudinal molecular studies, increasing these studies' cost-efficiency and statistical power for other diseases.
Our interdisciplinary approach begins with computational modelling of gene expression and symptom trajectories across multiple years following a patient's diagnosis. We will begin by combining molecular trajectories with baseline factors, including gender, ethnicity, and socio-economic status. Time-dependent changes in blood transcriptome and methylome will be associated with changes in patient symptoms, such as mean arterial pressure and 6 minute walk distance. Historically measured gene expression profiles will be used directly as longitudinal inputs in our machine learning models to predict future molecular profiles associated with pulmonary vascular remodelling and patient outcome.
To ensure scalability to many patient measures and time points, we will evaluate and extend a machine learning technique called Gaussian Processes, which has successfully modelled lower-dimensional longitudinal data such as body weights and CO2 emissions. The algorithm will be trained through an iterative process, by gradually increasing the amount of longitudinal data from patients' molecular profiles and electronic health records. Our multidisciplinary team of computer scientists, epidemiologists, biologists and clinicians, will enable the algorithm to use biologically relevant features and be scalable for clinical applications.
The data and methods from this project will not only provide PH patients with better information about their disease progression, but can also be used to understand other complex diseases and age-related disorders.
Publications
| Description | We have completed the first step towards the four aims in the project by generating longitudinal molecular data on pulmonary hypertension patients. In particular, we have measured DNA methylation across ~900,000 sites across the human genome and >1000 circulating proteins in the blood of these patients across four time points. Our initial quality assessment has not revealed technical artefacts that could compromise the data's quality. We hope this data (the largest of its kind for this disease) will promote scientific insights from our project and also from other studies. |
| Exploitation Route | We will make the two longitudinal datasets generated in this project accessible to all researchers. Currently, developers of computational methodologies have reached out to us about using the data to train and test their methods. Further use cases will emerge once we have published the dataset alongside our manuscript. |
| Sectors | Healthcare Pharmaceuticals and Medical Biotechnology |
| Description | GRADUATE PARTNER AGREEMENT |
| Amount | £35,390 (GBP) |
| Organisation | Health Data Research UK |
| Sector | Charity/Non Profit |
| Country | United Kingdom |
| Start | 09/2025 |
| End | 04/2026 |
| Description | Three-way collaboration between Imperial, Manchester and A*STAR on longitudinal machine learning of molecular and phenotypic trajectories of pulmonary hypertension. |
| Organisation | Agency for Science, Technology and Research (A*STAR) |
| Country | Singapore |
| Sector | Public |
| PI Contribution | As the Lead institution, Imperial College London and I provide the administrative and scientific backbone for the project. I serve as the Project Leader with overall management responsibility for the UKRI award , while the college handles all financial distributions and reporting to the funder. Scientifically, Imperial is responsible for the collection of clinical data and samples from its PH clinics and the National PAH Cohort , as well as performing the RNA and methylation profiling of blood samples. I specifically lead the data processing, quality control, and the iterative model and analyse phases of the research cycle to improve machine learning predictions of patient trajectories. |
| Collaborator Contribution | The Institute for Human Development and Potential (A*STAR) and the University of Manchester serve as the project's primary computational hubs, transforming raw data into predictive insights. A*STAR, led by Dr. Varsha Gupta, manages a secure cloud platform dedicated to federated data analysis. This infrastructure allows for the compilation of complex longitudinal trajectories from molecular and symptom data while maintaining security. Building on this, the University of Manchester, led by Dr. Mauricio Alvarez Lopez, applies advanced machine learning models to integrate these high-dimensional trajectories. Together, they form an iterative feedback loop where their computational results are interpreted for clinical plausibility by the Lead, which then informs the next phase of sample collection and model refinement. |
| Impact | - longitudinal measures at 4 timepoints of whole methylome of 210 patients with different combinations of pulmonary hypertension and comorbidities - longitudinal measures at 4 timepoints of proteome of 1000 proteins in 210 patients with different combinations of pulmonary hypertension and comorbidities - extracted cardiovascular function measures and clinical data across timepoints for the 210 patients |
| Start Year | 2025 |
| Description | Three-way collaboration between Imperial, Manchester and A*STAR on longitudinal machine learning of molecular and phenotypic trajectories of pulmonary hypertension. |
| Organisation | University of Manchester |
| Country | United Kingdom |
| Sector | Academic/University |
| PI Contribution | As the Lead institution, Imperial College London and I provide the administrative and scientific backbone for the project. I serve as the Project Leader with overall management responsibility for the UKRI award , while the college handles all financial distributions and reporting to the funder. Scientifically, Imperial is responsible for the collection of clinical data and samples from its PH clinics and the National PAH Cohort , as well as performing the RNA and methylation profiling of blood samples. I specifically lead the data processing, quality control, and the iterative model and analyse phases of the research cycle to improve machine learning predictions of patient trajectories. |
| Collaborator Contribution | The Institute for Human Development and Potential (A*STAR) and the University of Manchester serve as the project's primary computational hubs, transforming raw data into predictive insights. A*STAR, led by Dr. Varsha Gupta, manages a secure cloud platform dedicated to federated data analysis. This infrastructure allows for the compilation of complex longitudinal trajectories from molecular and symptom data while maintaining security. Building on this, the University of Manchester, led by Dr. Mauricio Alvarez Lopez, applies advanced machine learning models to integrate these high-dimensional trajectories. Together, they form an iterative feedback loop where their computational results are interpreted for clinical plausibility by the Lead, which then informs the next phase of sample collection and model refinement. |
| Impact | - longitudinal measures at 4 timepoints of whole methylome of 210 patients with different combinations of pulmonary hypertension and comorbidities - longitudinal measures at 4 timepoints of proteome of 1000 proteins in 210 patients with different combinations of pulmonary hypertension and comorbidities - extracted cardiovascular function measures and clinical data across timepoints for the 210 patients |
| Start Year | 2025 |
