Multicellular factor analysis of single-cell data for a tissue-centric understanding of disease

Abstract

Biomedical single-cell atlases describe disease at the cellular level. However, analysis of this data commonly focuses on cell-type centric pairwise cross-condition comparisons, disregarding the multicellular nature of disease processes. Here we propose multicellular factor analysis for the unsupervised analysis of samples from cross-condition single-cell atlases and the identification of multicellular programs associated with disease. Our strategy, which repurposes group factor analysis as implemented in multi-omics factor analysis, incorporates the variation of patient samples across cell-types or other tissue-centric features, such as cell compositions or spatial relationships, and enables the joint analysis of multiple patient cohorts, facilitating the integration of atlases. We applied our framework to a collection of acute and chronic human heart failure atlases and described multicellular processes of cardiac remodeling, independent to cellular compositions and their local organization, that were conserved in independent spatial and bulk transcriptomics datasets. In sum, our framework serves as an exploratory tool for unsupervised analysis of cross-condition single-cell atlases and allows for the integration of the measurements of patient cohorts across distinct data modalities.

Data availability

The datasets and computer code produced in this study are available in the following databases:-All scripts related to this manuscript can be consulted here: https://github.com/saezlab/MOFAcell.-The R package implementing multicellular factor analysis can be found in:https://github.com/saezlab/MOFAcellulaR-The python implementation of multicellular factor analysis is available here:https://liana-py.readthedocs.io/en/latest/notebooks/mofacellular.html-A Zenodo entry containing data associated to this manuscript can be accessed here: https://zenodo.org/record/8082895.

The following previously published data sets were used

Article and author information

Author details

  1. Ricardo Omar Ramirez Flores

    Faculty of Medicine, Heidelberg University, Heidelberg, Germany
    For correspondence
    roramirezf@uni-heidelberg.de
    Competing interests
    No competing interests declared.
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0003-0087-371X
  2. Jan David Lanzer

    Faculty of Medicine, Heidelberg University, Heidelberg, Germany
    Competing interests
    No competing interests declared.
  3. Daniel Dimitrov

    Faculty of Medicine, Heidelberg University, Heidelberg, Germany
    Competing interests
    No competing interests declared.
  4. Britta Velten

    Centre for Organismal Studies, Heidelberg University, Heidelberg, Germany
    Competing interests
    No competing interests declared.
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0002-8397-3515
  5. Julio Saez-Rodriguez

    Faculty of Medicine, Heidelberg University, Heidelberg, Germany
    For correspondence
    pub.saez@uni-heidelberg.de
    Competing interests
    Julio Saez-Rodriguez, reports funding from GSK, Pfizer and Sanofi and fees/honoraria from Travere Therapeutics, Stadapharm, Astex, Pfizer and Grunenthal..
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0002-8552-8976

Funding

DFG CRC 1550 (464424253)

  • Ricardo Omar Ramirez Flores
  • Julio Saez-Rodriguez

Informatics for Life

  • Jan David Lanzer
  • Julio Saez-Rodriguez

EU ITN Marie Curie StrategyCKD (860329)

  • Daniel Dimitrov
  • Julio Saez-Rodriguez

The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication.

Reviewing Editor

  1. Jihwan Park, Gwangju Institute of Science and Technology, Republic of Korea

Version history

  1. Preprint posted: February 23, 2023 (view preprint)
  2. Received: October 5, 2023
  3. Accepted: November 14, 2023
  4. Accepted Manuscript published: November 22, 2023 (version 1)
  5. Version of Record published: December 13, 2023 (version 2)

Copyright

© 2023, Ramirez Flores et al.

This article is distributed under the terms of the Creative Commons Attribution License permitting unrestricted use and redistribution provided that the original author and source are credited.

Metrics

  • 2,727
    views
  • 419
    downloads
  • 6
    citations

Views, downloads and citations are aggregated across all versions of this paper published by eLife.

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Ricardo Omar Ramirez Flores
  2. Jan David Lanzer
  3. Daniel Dimitrov
  4. Britta Velten
  5. Julio Saez-Rodriguez
(2023)
Multicellular factor analysis of single-cell data for a tissue-centric understanding of disease
eLife 12:e93161.
https://doi.org/10.7554/eLife.93161

Share this article

https://doi.org/10.7554/eLife.93161

Further reading

    1. Computational and Systems Biology
    2. Neuroscience
    Andrea I Luppi, Pedro AM Mediano ... Emmanuel A Stamatakis
    Research Article

    How is the information-processing architecture of the human brain organised, and how does its organisation support consciousness? Here, we combine network science and a rigorous information-theoretic notion of synergy to delineate a ‘synergistic global workspace’, comprising gateway regions that gather synergistic information from specialised modules across the human brain. This information is then integrated within the workspace and widely distributed via broadcaster regions. Through functional MRI analysis, we show that gateway regions of the synergistic workspace correspond to the human brain’s default mode network, whereas broadcasters coincide with the executive control network. We find that loss of consciousness due to general anaesthesia or disorders of consciousness corresponds to diminished ability of the synergistic workspace to integrate information, which is restored upon recovery. Thus, loss of consciousness coincides with a breakdown of information integration within the synergistic workspace of the human brain. This work contributes to conceptual and empirical reconciliation between two prominent scientific theories of consciousness, the Global Neuronal Workspace and Integrated Information Theory, while also advancing our understanding of how the human brain supports consciousness through the synergistic integration of information.

    1. Computational and Systems Biology
    2. Genetics and Genomics
    Ardalan Naseri, Degui Zhi, Shaojie Zhang
    Research Article Updated

    Runs-of-homozygosity (ROH) segments, contiguous homozygous regions in a genome were traditionally linked to families and inbred populations. However, a growing literature suggests that ROHs are ubiquitous in outbred populations. Still, most existing genetic studies of ROH in populations are limited to aggregated ROH content across the genome, which does not offer the resolution for mapping causal loci. This limitation is mainly due to a lack of methods for the efficient identification of shared ROH diplotypes. Here, we present a new method, ROH-DICE (runs-of-homozygous diplotype cluster enumerator), to find large ROH diplotype clusters, sufficiently long ROHs shared by a sufficient number of individuals, in large cohorts. ROH-DICE identified over 1 million ROH diplotypes that span over 100 single nucleotide polymorphisms (SNPs) and are shared by more than 100 UK Biobank participants. Moreover, we found significant associations of clustered ROH diplotypes across the genome with various self-reported diseases, with the strongest associations found between the extended human leukocyte antigen (HLA) region and autoimmune disorders. We found an association between a diplotype covering the homeostatic iron regulator (HFE) gene and hemochromatosis, even though the well-known causal SNP was not directly genotyped or imputed. Using a genome-wide scan, we identified a putative association between carriers of an ROH diplotype in chromosome 4 and an increase in mortality among COVID-19 patients (p-value = 1.82 × 10−11). In summary, our ROH-DICE method, by calling out large ROH diplotypes in a large outbred population, enables further population genetics into the demographic history of large populations. More importantly, our method enables a new genome-wide mapping approach for finding disease-causing loci with multi-marker recessive effects at a population scale.