A phylogenetic transform enhances analysis of compositional microbiota data
Abstract
Surveys of microbial communities (microbiota), typically measured as relative abundance of species, have illustrated the importance of these communities in human health and disease. Yet, statistical artifacts commonly plague the analysis of relative abundance data. Here, we introduce the PhILR transform, which incorporates microbial evolutionary models with the isometric log-ratio transform to allow off-the-shelf statistical tools to be safely applied to microbiota surveys. We demonstrate that analyses of community-level structure can be applied to PhILR transformed data with performance on benchmarks rivaling or surpassing standard tools. Additionally, By decomposing distance in the PhILR transformed space, we identified neighboring clades that may have adapted to distinct human body sites. Decomposing variance revealed that covariation of bacterial clades within human body sites increases with phylogenetic relatedness. Together, these findings illustrate how the PhILR transform combines statistical and phylogenetic models to overcome compositional data challenges and enable evolutionary insights relevant to microbial communities.
Data availability
-
Human Microbiome ProjectPublicly available at HMPDACC (v35 download of files 6, 9, and 10).
-
Costello Skin SitesPublicly available as part of the FEMS Benchmark dataset (2011) provided Dan Knights.
-
Global PatternsPublicly available and provided as part of the phyloseq R package as 'GlobalPatterns'.
Article and author information
Author details
Funding
Global Probiotics Council (Young Investigator Grant for Probiotics Research)
- Lawrence A David
Searle Scholars Program (15-SSP-184 Research Agreement)
- Lawrence A David
Alfred P. Sloan Foundation (BR2014-003)
- Lawrence A David
The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication.
Reviewing Editor
- Anthony Fodor, University of North Carolina at Charlotte
Version history
- Received: September 27, 2016
- Accepted: February 13, 2017
- Accepted Manuscript published: February 15, 2017 (version 1)
- Version of Record published: February 27, 2017 (version 2)
Copyright
© 2017, Silverman et al.
This article is distributed under the terms of the Creative Commons Attribution License permitting unrestricted use and redistribution provided that the original author and source are credited.
Metrics
-
- 11,403
- views
-
- 1,717
- downloads
-
- 236
- citations
Views, downloads and citations are aggregated across all versions of this paper published by eLife.
Download links
Downloads (link to download the article as PDF)
Open citations (links to open the citations from this article in various online reference manager services)
Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)
Further reading
-
- Genetics and Genomics
PARP-1 is central to transcriptional regulation under both normal and stress conditions, with the governing mechanisms yet to be fully understood. Our biochemical and ChIP-seq-based analyses showed that PARP-1 binds specifically to active histone marks, particularly H4K20me1. We found that H4K20me1 plays a critical role in facilitating PARP-1 binding and the regulation of PARP-1-dependent loci during both development and heat shock stress. Here, we report that the sole H4K20 mono-methylase, pr-set7, and parp-1 Drosophila mutants undergo developmental arrest. RNA-seq analysis showed an absolute correlation between PR-SET7- and PARP-1-dependent loci expression, confirming co-regulation during developmental phases. PARP-1 and PR-SET7 are both essential for activating hsp70 and other heat shock genes during heat stress, with a notable increase of H4K20me1 at their gene body. Mutating pr-set7 disrupts monomethylation of H4K20 along heat shock loci and abolish PARP-1 binding there. These data strongly suggest that H4 monomethylation is a key triggering point in PARP-1 dependent processes in chromatin.
-
- Cancer Biology
- Genetics and Genomics
Enhancers are critical for regulating tissue-specific gene expression, and genetic variants within enhancer regions have been suggested to contribute to various cancer-related processes, including therapeutic resistance. However, the precise mechanisms remain elusive. Using a well-defined drug-gene pair, we identified an enhancer region for dihydropyrimidine dehydrogenase (DPD, DPYD gene) expression that is relevant to the metabolism of the anti-cancer drug 5-fluorouracil (5-FU). Using reporter systems, CRISPR genome-edited cell models, and human liver specimens, we demonstrated in vitro and vivo that genotype status for the common germline variant (rs4294451; 27% global minor allele frequency) located within this novel enhancer controls DPYD transcription and alters resistance to 5-FU. The variant genotype increases recruitment of the transcription factor CEBPB to the enhancer and alters the level of direct interactions between the enhancer and DPYD promoter. Our data provide insight into the regulatory mechanisms controlling sensitivity and resistance to 5-FU.