Projects Offered

Dorothee Dormann  René Ketting  Edward Lemke  Andreas Walther  Eva Wolf  Shuqing Xu  Johannes Mayer_Activate  Johannes Mayer_DCTraining  Vincent ten Cate 

Causal multi-omics modeling of disease biology

1 PhD project offered in the IPP winter call Molecular Biomedicine & Ageing

Scientific background

Many existing multi-omics methods focus on dimension reduction or joint latent representations to integrate heterogeneous molecular data, but are often less explicit about biological directionality and causal inference. This project instead builds on the structure of the central dogma, using genetic variation as an anchor for causal inference across molecular layers (genomics, transcriptomics, proteomics, and metabolomics), enabling more principled identification of disease-relevant proteins and pathways. A key additional aspect is the use of functional structure within the proteome—such as protein–protein interaction patterns, structural similarity, and data-driven embeddings derived from sequence or genetic perturbation—as prior information, complementing or extending curated pathway knowledge. This combination allows us to move beyond purely associative integration toward structure-aware, causally interpretable models of molecular disease mechanisms, with a focus on identifying proteins and pathways involved in cardiovascular disease.

PhD project: Causal multi-omics modeling of disease biology

We are seeking a PhD candidate to join our young team (Computational Systems Medicine) within the context of the BMFTR-funded DIASyM (https://diasym.mscoresys.de/) project, to develop new statistical and computational methods for integrating multi-omics data to infer causal mechanisms of disease. We are a new group comprised of a molecular epidemiologist (group leader) and applied bioinformaticians, and are currently looking for a methodologist to complement our team.


The project focuses on building principled models that combine genetic association data (GWAS), molecular QTLs (eQTLs and pQTLs), (tissue-specific) transcriptomics (e.g. GTEx), and proteomics, with additional layers such as lipidomics and metabolomics where available. Access to large datasets is guaranteed, from local large cohort studies with multiomics phenotyping to external datasets like UK Biobank. The central goal of the project is to perform causal inference of disease-relevant proteins and molecular pathways. The disease area of interest is cardiovascular disease.
 

A possible direction for the work is the development of probabilistic graphical models that represent proteins as latent causal drivers of disease, while treating other omics layers as noisy, partially mediated observations. These models will incorporate biologically informed structure, including protein–protein interaction networks derived from data-driven sources such as protein embeddings, genetic variation (e.g. pQTLs) and experimental perturbation data, rather than relying on curated pathway databases (e.g. KEGG, Reactome) that introduce a human bias.
 

The successful candidate may incorporate methods that integrate:
- Mendelian randomization and genetic instruments
- Bayesian hierarchical models and Gaussian graphical models
- Multi-layer data integration across tissues and omics modalities
- Network-based regularization informed by protein structure and embeddings
- Scalable inference methods for genome-wide applications
 

However, own ideas on how to approach the project are highly welcomed. The project sits at the intersection of statistical genetics, systems biology, and machine learning, with strong emphasis on methodological development.   
 

Tasks of the PhD Student
- Develop and evaluate statistical and machine learning models
- Publish results in peer-reviewed journals


Desired Qualifications
- Master’s degree in statistics, mathematics, data science, bioinformatics, physics, engineering, or a related quantitative field
- Knowledge of statistics 
- Interest in machine learning
- Experience with R and/or Python
- Good English skills
 

If you are interested in this project, please select ten Cate your group preference in the IPP application platform.

 

Publications relevant to the project

Ingold M, Müller C, Krolevets M, Yapıcı E, Rapp S, Schuster AK, Tesarz J, Heinrich I, Weinmann-Menke J, Strauch K, Lackner KJ, Konstantinides S, Ruf W, Andrade-Navarro MA, Niehrs C, Gori T, Lurz P, Wild PS, Ten Cate V; GHS Research Consortium (2026) Blood DNA Methylation Patterns Across Carotid, Coronary, and Peripheral Atherosclerosis: A Comparative Analysis in 2 Prospective Cohorts. JACC. Jul 21;88(3):302-320. Link

Han J, Ten Cate V, Li-Gao R, Robles AP, Raaja Sulochana H, Andrade-Navarro MA, Ao L, Noordam R, Martinez-Perez A, Sabater-Lleal M, Soria JM, Souto JC, Rosendaal FR, Wild PS, van Hylckama Vlieg A (2026) Genome-wide identification of loci associated with plasma coagulation factor IX activity. J Thromb Haemost. Feb;24(2):545-557. Link

Schmitt F, Ten Cate V, Fischer Z, Hagen M, Steigenberger BA, Tenzer S, Wild PS, Schmidlin T (2025) Metabolic Profiling of the EmDia Cohort by LC-MS Reveals Empagliflozin-Intake Associated Regulation of 1,5-anhydroglucitol and Urate. Proteomics. Jan;26(1):44-56. doi: 10.1002/pmic.70075. Link

Anyaegbunam U, ten Cate V, Bauer K, Schmidlin T, Distler U, Tenzer S, Araldi E, Bindila L, Wild PS, Andrade-Navarro M (2026) Integrative Multi-Omics Analysis Reveals Lipid/Metabolite Dysregulation and Temporal Decoupling in Disease Progression. Sci Rep. (accepted). 

 

Contact details

Jun-Prof. Dr. Vincent ten Cate
Email
Website