October 1, 2026
Project overview
This project investigates whether identifiable variational autoencoders (iVAEs) can support causal discovery in microbiome data when unobserved factors confound relationships with health or well-being. Lusis has slightly more than 2,000 questionnaires paired with microbiome analyses. These cross-sectional observations provide the context for a 200-hour group project in 2026–2027.
Students will learn a compact microbiome representation using questionnaire variables as auxiliary information [1, 2], then apply Fast Causal Inference (FCI) to the estimated factors and selected health and context variables [6]. FCI accommodates hidden variables and returns a partial ancestral graph representing compatible causal structures. The central question is whether identifiable components provide suitable causal variables: the conditional independence assumptions of standard iVAE do not automatically establish a biological causal graph [3].
A simulated causal model with known structure and hidden confounding will evaluate the complete pipeline. Comparing FCI on true factors with FCI on learned factors will help distinguish graph-estimation errors from representation failures. Experiments will vary the informativeness of auxiliary variables and assess factor recovery, identifiable graph relations, incorrect orientations and stability. Principal component analysis provides a simple baseline. Microbial composition, zeros and technical batches will be handled explicitly [4, 5]; posterior-collapse diagnostics and CI-iVAE provide possible supporting controls [7].
The pipeline will then be explored on Lusis data, with separate development and evaluation patients. The decoder will connect selected stable factors to microbial abundance contrasts. These remain model-based hypotheses, requiring further observations or experiments for validation. Expected deliverables include reproducible code, a simulation benchmark, an annotated exploratory graph, a scientific report and a presentation. Multimodal causal representation learning offers a possible subsequent extension [8].
References
[1] Khemakhem, I., Kingma, D. P., Monti, R. P., & Hyvärinen, A. (2020). Variational Autoencoders and Nonlinear ICA: A Unifying Framework. AISTATS, PMLR 108, 2207–2217. https://proceedings.mlr.press/v108/khemakhem20a.html
[2] Xi, Q., & Bloem-Reddy, B. (2023; preprint 2022). Indeterminacy in Generative Models: Characterization and Strong Identifiability. AISTATS. https://arxiv.org/abs/2206.00801
[3] Yao, D., Rancati, D., Cadei, R., Fumero, M., & Locatello, F. (2025; preprint 2024). Unifying Causal Representation Learning with the Invariance Principle. ICLR. https://arxiv.org/abs/2409.02772
[4] Gloor, G. B., Macklaim, J. M., Pawlowsky-Glahn, V., & Egozcue, J. J. (2017). Microbiome Datasets Are Compositional: And This Is Not Optional. Frontiers in Microbiology, 8, 2224. https://pmc.ncbi.nlm.nih.gov/articles/PMC5695134/
[5] Vujkovic-Cvijin, I., et al. (2020). Host variables confound gut microbiota studies of human disease. Nature, 587, 448–454. https://www.nature.com/articles/s41586-020-2881-9
[6] Colombo, D., Maathuis, M. H., Kalisch, M., & Richardson, T. S. (2012). Learning high-dimensional directed acyclic graphs with latent and selection variables. The Annals of Statistics, 40(1), 294–321. https://arxiv.org/abs/1104.5617
[7] Kim, Y.-G., Liu, Y., & Wei, X.-X. (2023). Covariate-informed Representation Learning to Prevent Posterior Collapse of iVAE. AISTATS, PMLR 206, 2641–2660. https://proceedings.mlr.press/v206/kim23c.html
[8] Sun, Y., et al. (2024, revised 2025). Causal Representation Learning from Multimodal Biomedical Observations. https://arxiv.org/abs/2411.06518