Causal discovery in microbiome data with identifiable variational autoencoders

October 1, 2026

Project overview

This project investigates whether identifiable variational autoencoders (iVAEs) can support causal discovery in microbiome data when unobserved factors confound relationships with health or well-being. Lusis has slightly more than 2,000 questionnaires paired with microbiome analyses. These cross-sectional observations provide the context for a 200-hour group project in 2026–2027.

Students will learn a compact microbiome representation using questionnaire variables as auxiliary information [1, 2], then apply Fast Causal Inference (FCI) to the estimated factors and selected health and context variables [6]. FCI accommodates hidden variables and returns a partial ancestral graph representing compatible causal structures. The central question is whether identifiable components provide suitable causal variables: the conditional independence assumptions of standard iVAE do not automatically establish a biological causal graph [3].

A simulated causal model with known structure and hidden confounding will evaluate the complete pipeline. Comparing FCI on true factors with FCI on learned factors will help distinguish graph-estimation errors from representation failures. Experiments will vary the informativeness of auxiliary variables and assess factor recovery, identifiable graph relations, incorrect orientations and stability. Principal component analysis provides a simple baseline. Microbial composition, zeros and technical batches will be handled explicitly [4, 5]; posterior-collapse diagnostics and CI-iVAE provide possible supporting controls [7].

The pipeline will then be explored on Lusis data, with separate development and evaluation patients. The decoder will connect selected stable factors to microbial abundance contrasts. These remain model-based hypotheses, requiring further observations or experiments for validation. Expected deliverables include reproducible code, a simulation benchmark, an annotated exploratory graph, a scientific report and a presentation. Multimodal causal representation learning offers a possible subsequent extension [8].

References

[1] Khemakhem, I., Kingma, D. P., Monti, R. P., & Hyvärinen, A. (2020). Variational Autoencoders and Nonlinear ICA: A Unifying Framework. AISTATS, PMLR 108, 2207–2217. https://proceedings.mlr.press/v108/khemakhem20a.html

[2] Xi, Q., & Bloem-Reddy, B. (2023; preprint 2022). Indeterminacy in Generative Models: Characterization and Strong Identifiability. AISTATS. https://arxiv.org/abs/2206.00801

[3] Yao, D., Rancati, D., Cadei, R., Fumero, M., & Locatello, F. (2025; preprint 2024). Unifying Causal Representation Learning with the Invariance Principle. ICLR. https://arxiv.org/abs/2409.02772

[4] Gloor, G. B., Macklaim, J. M., Pawlowsky-Glahn, V., & Egozcue, J. J. (2017). Microbiome Datasets Are Compositional: And This Is Not Optional. Frontiers in Microbiology, 8, 2224. https://pmc.ncbi.nlm.nih.gov/articles/PMC5695134/

[5] Vujkovic-Cvijin, I., et al. (2020). Host variables confound gut microbiota studies of human disease. Nature, 587, 448–454. https://www.nature.com/articles/s41586-020-2881-9

[6] Colombo, D., Maathuis, M. H., Kalisch, M., & Richardson, T. S. (2012). Learning high-dimensional directed acyclic graphs with latent and selection variables. The Annals of Statistics, 40(1), 294–321. https://arxiv.org/abs/1104.5617

[7] Kim, Y.-G., Liu, Y., & Wei, X.-X. (2023). Covariate-informed Representation Learning to Prevent Posterior Collapse of iVAE. AISTATS, PMLR 206, 2641–2660. https://proceedings.mlr.press/v206/kim23c.html

[8] Sun, Y., et al. (2024, revised 2025). Causal Representation Learning from Multimodal Biomedical Observations. https://arxiv.org/abs/2411.06518