Mosaic cis-regulatory evolution drives transcriptional partitioning of HERVH endogenous retrovirus in the human embryo
Abstract
The human endogenous retrovirus type-H (HERVH) family is expressed in the preimplantation embryo. A subset of these elements are specifically transcribed in pluripotent stem cells where they appear to exert regulatory activities promoting self-renewal and pluripotency. How HERVH elements achieve such transcriptional specificity remains poorly understood. To uncover the sequence features underlying HERVH transcriptional activity, we performed a phyloregulatory analysis of the long terminal repeats (LTR7) of the HERVH family, which harbor its promoter, using a wealth of regulatory genomics data. We found that the family includes at least 8 previously unrecognized subfamilies that have been active at different timepoints in primate evolution and display distinct expression patterns during human embryonic development. Notably, nearly all HERVH elements transcribed in ESCs belong to one of the youngest subfamilies we dubbed LTR7up. LTR7 sequence evolution was driven by a mixture of mutational processes, including point mutations, duplications, and multiple recombination events between subfamilies, that led to transcription factor binding motif modules characteristic of each subfamily. Using a reporter assay, we show that one such motif, a predicted SOX2/3 binding site unique to LTR7up, is essential for robust promoter activity in induced pluripotent stem cells. Together these findings illuminate the mechanisms by which HERVH diversified its expression pattern during evolution to colonize distinct cellular niches within the human embryo.
Data availability
Scripts, data tables, and notes for figures 1-4,6a and supplemental figures 1-1,2-1,3-1,4-1,5-1,6-2 by TAC and JDC - https://github.com/LumpLord/Mosaic-cis-regulatory-evolution-drives-transcriptional-partitioning-of-HERVH-endogenous-retrovirus..Scripts and data tables by MS for figures 5,6c and supplemental figures 6-1,6-3,5-2 - https://github.com/Manu-1512/LTR7-up
-
Transcription factor binding dynamics during human ES cell differentiationNCBI Gene Expression Omnibus, GSE61475.
-
3D Chromosome Regulatory Landscape of Human Pluripotent CellsNCBI Gene Expression Omnibus, GSE69647.
-
ChIP-exo of human KRAB-ZNFs transduced in HEK 293T cells and KAP1 in hES H1 cellsNCBI Gene Expression Omnibus, GSE78099.
-
Repeat elements study in pluripotent stem cellsNCBI Gene Expression Omnibus, GSE54726.
-
Principles of Signalling Pathway Modulation for Enhancing Human Naïve Pluripotency Induction [ChIP-seq]NCBI Gene Expression Omnibus, GSE125553.
-
Tracing pluripotency of human early embryos and embryonic stem cells by single cell RNA-seqNCBI Gene Expression Omnibus, GSE36552.
-
Single-Cell RNA-seq Defines the Three Cell Lineages of the Human BlastocystNCBI Gene Expression Omnibus, GSE66507.
Article and author information
Author details
Funding
National Institutes of Health (GM112972)
- Cédric Feschotte
National Institutes of Health (HG009391)
- Cédric Feschotte
National Institutes of Health (GM122550)
- Cédric Feschotte
Cornell Center for Vertebrate Genomics
- Thomas Carter
Howard Hughes Medical Institute
- John L Rinn
National Institutes of Health (GM099117)
- John L Rinn
Cornell Presidential Fellow Program
- Manvendra Singh
The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication.
Copyright
© 2022, Carter et al.
This article is distributed under the terms of the Creative Commons Attribution License permitting unrestricted use and redistribution provided that the original author and source are credited.
Metrics
-
- 3,893
- views
-
- 457
- downloads
-
- 38
- citations
Views, downloads and citations are aggregated across all versions of this paper published by eLife.
Download links
Downloads (link to download the article as PDF)
Open citations (links to open the citations from this article in various online reference manager services)
Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)
Further reading
-
- Genetics and Genomics
Flavin-containing monooxygenases (FMOs) are a conserved family of xenobiotic enzymes upregulated in multiple longevity interventions, including nematode and mouse models. Previous work supports that C. elegans fmo-2 promotes longevity, stress resistance, and healthspan by rewiring endogenous metabolism. However, there are five C. elegans FMOs and five mammalian FMOs, and it is not known whether promoting longevity and health benefits is a conserved role of this gene family. Here, we report that expression of C. elegans fmo-4 promotes lifespan extension and paraquat stress resistance downstream of both dietary restriction and inhibition of mTOR. We find that overexpression of fmo-4 in just the hypodermis is sufficient for these benefits, and that this expression significantly modifies the transcriptome. By analyzing changes in gene expression, we find that genes related to calcium signaling are significantly altered downstream of fmo-4 expression. Highlighting the importance of calcium homeostasis in this pathway, fmo-4 overexpressing animals are sensitive to thapsigargin, an ER stressor that inhibits calcium flux from the cytosol to the ER lumen. This calcium/fmo-4 interaction is solidified by data showing that modulating intracellular calcium with either small molecules or genetics can change expression of fmo-4 and/or interact with fmo-4 to affect lifespan and stress resistance. Further analysis supports a pathway where fmo-4 modulates calcium homeostasis downstream of activating transcription factor-6 (atf-6), whose knockdown induces and requires fmo-4 expression. Together, our data identify fmo-4 as a longevity-promoting gene whose actions interact with known longevity pathways and calcium homeostasis.
-
- Genetics and Genomics
One of the goals of synthetic biology is to enable the design of arbitrary molecular circuits with programmable inputs and outputs. Such circuits bridge the properties of electronic and natural circuits, processing information in a predictable manner within living cells. Genome editing is a potentially powerful component of synthetic molecular circuits, whether for modulating the expression of a target gene or for stably recording information to genomic DNA. However, programming molecular events such as protein-protein interactions or induced proximity as triggers for genome editing remains challenging. Here, we demonstrate a strategy termed ‘P3 editing’, which links protein-protein proximity to the formation of a functional CRISPR-Cas9 dual-component guide RNA. By engineering the crRNA:tracrRNA interaction, we demonstrate that various known protein-protein interactions, as well as the chemically induced dimerization of protein domains, can be used to activate prime editing or base editing in human cells. Additionally, we explore how P3 editing can incorporate outputs from ADAR-based RNA sensors, potentially allowing specific RNAs to induce specific genome edits within a larger circuit. Our strategy enhances the controllability of CRISPR-based genome editing, facilitating its use in synthetic molecular circuits deployed in living cells.