Projects Offered
Dorothee Dormann René Ketting Edward Lemke Andreas Walther Eva Wolf Shuqing Xu Johannes Mayer_Activate Johannes Mayer_DCTraining Vincent ten CateCausal multi-omics modeling of disease biology
1 PhD project offered in the IPP winter call Molecular Biomedicine & Ageing
Scientific background
Many existing multi-omics methods focus on dimension reduction or joint latent representations to integrate heterogeneous molecular data, but are often less explicit about biological directionality and causal inference. This project instead builds on the structure of the central dogma, using genetic variation as an anchor for causal inference across molecular layers (genomics, transcriptomics, proteomics, and metabolomics), enabling more principled identification of disease-relevant proteins and pathways. A key additional aspect is the use of functional structure within the proteome—such as protein–protein interaction patterns, structural similarity, and data-driven embeddings derived from sequence or genetic perturbation—as prior information, complementing or extending curated pathway knowledge. This combination allows us to move beyond purely associative integration toward structure-aware, causally interpretable models of molecular disease mechanisms, with a focus on identifying proteins and pathways involved in cardiovascular disease.
PhD project: Causal multi-omics modeling of disease biology
We are seeking a PhD candidate to join our young team (Computational Systems Medicine) within the context of the BMFTR-funded DIASyM (https://diasym.mscoresys.de/) project, to develop new statistical and computational methods for integrating multi-omics data to infer causal mechanisms of disease. We are a new group comprised of a molecular epidemiologist (group leader) and applied bioinformaticians, and are currently looking for a methodologist to complement our team.
The project focuses on building principled models that combine genetic association data (GWAS), molecular QTLs (eQTLs and pQTLs), (tissue-specific) transcriptomics (e.g. GTEx), and proteomics, with additional layers such as lipidomics and metabolomics where available. Access to large datasets is guaranteed, from local large cohort studies with multiomics phenotyping to external datasets like UK Biobank. The central goal of the project is to perform causal inference of disease-relevant proteins and molecular pathways. The disease area of interest is cardiovascular disease.
A possible direction for the work is the development of probabilistic graphical models that represent proteins as latent causal drivers of disease, while treating other omics layers as noisy, partially mediated observations. These models will incorporate biologically informed structure, including protein–protein interaction networks derived from data-driven sources such as protein embeddings, genetic variation (e.g. pQTLs) and experimental perturbation data, rather than relying on curated pathway databases (e.g. KEGG, Reactome) that introduce a human bias.
The successful candidate may incorporate methods that integrate:
- Mendelian randomization and genetic instruments
- Bayesian hierarchical models and Gaussian graphical models
- Multi-layer data integration across tissues and omics modalities
- Network-based regularization informed by protein structure and embeddings
- Scalable inference methods for genome-wide applications
However, own ideas on how to approach the project are highly welcomed. The project sits at the intersection of statistical genetics, systems biology, and machine learning, with strong emphasis on methodological development.
Tasks of the PhD Student
- Develop and evaluate statistical and machine learning models
- Publish results in peer-reviewed journals
Desired Qualifications
- Master’s degree in statistics, mathematics, data science, bioinformatics, physics, engineering, or a related quantitative field
- Knowledge of statistics
- Interest in machine learning
- Experience with R and/or Python
- Good English skills
If you are interested in this project, please select ten Cate your group preference in the IPP application platform.
Publications relevant to the project
Ingold M, Müller C, Krolevets M, Yapıcı E, Rapp S, Schuster AK, Tesarz J, Heinrich I, Weinmann-Menke J, Strauch K, Lackner KJ, Konstantinides S, Ruf W, Andrade-Navarro MA, Niehrs C, Gori T, Lurz P, Wild PS, Ten Cate V; GHS Research Consortium (2026) Blood DNA Methylation Patterns Across Carotid, Coronary, and Peripheral Atherosclerosis: A Comparative Analysis in 2 Prospective Cohorts. JACC. Jul 21;88(3):302-320. Link
Han J, Ten Cate V, Li-Gao R, Robles AP, Raaja Sulochana H, Andrade-Navarro MA, Ao L, Noordam R, Martinez-Perez A, Sabater-Lleal M, Soria JM, Souto JC, Rosendaal FR, Wild PS, van Hylckama Vlieg A (2026) Genome-wide identification of loci associated with plasma coagulation factor IX activity. J Thromb Haemost. Feb;24(2):545-557. Link
Schmitt F, Ten Cate V, Fischer Z, Hagen M, Steigenberger BA, Tenzer S, Wild PS, Schmidlin T (2025) Metabolic Profiling of the EmDia Cohort by LC-MS Reveals Empagliflozin-Intake Associated Regulation of 1,5-anhydroglucitol and Urate. Proteomics. Jan;26(1):44-56. doi: 10.1002/pmic.70075. Link
Anyaegbunam U, ten Cate V, Bauer K, Schmidlin T, Distler U, Tenzer S, Araldi E, Bindila L, Wild PS, Andrade-Navarro M (2026) Integrative Multi-Omics Analysis Reveals Lipid/Metabolite Dysregulation and Temporal Decoupling in Disease Progression. Sci Rep. (accepted).
