Systematic analysis of naturally occurring insertions and deletions that alter transcription factor spacing identifies tolerant and sensitive transcription factor pairs

  1. Zeyang Shen
  2. Rick Z Li
  3. Thomas A Prohaska
  4. Marten A Hoeksema
  5. Nathan J Spann
  6. Jenhan Tao
  7. Gregory J Fonseca
  8. Thomas Le
  9. Lindsey K Stolze
  10. Mashito Sakai
  11. Casey E Romanoski
  12. Christopher K Glass  Is a corresponding author
  1. University of California, San Diego, United States
  2. University of Arizona, United States
  3. Nippon Medical School, Japan
  4. University of California San Diego, United States

Abstract

Regulation of gene expression requires the combinatorial binding of sequence-specific transcription factors (TFs) at promoters and enhancers. Prior studies showed that alterations in the spacing between TF binding sites can influence promoter and enhancer activity. However, the relative importance of TF spacing alterations resulting from naturally occurring insertions and deletions (InDels) has not been systematically analyzed. To address this question, we first characterized the genome-wide spacing relationships of 73 TFs in human K562 cells as determined by ChIP-seq. We found a dominant pattern of a relaxed range of spacing between collaborative factors, including 45 TFs exclusively exhibiting relaxed spacing with their binding partners. Next, we exploited millions of InDels provided by genetically diverse mouse strains and human individuals to investigate the effects of altered spacing on TF binding and local histone acetylation. These analyses suggested that spacing alterations resulting from naturally occurring InDels are generally tolerated in comparison to genetic variants directly affecting TF binding sites. To experimentally validate this prediction, we introduced synthetic spacing alterations between PU.1 and C/EBPβ binding sites at six endogenous genomic loci in a macrophage cell line. Remarkably, collaborative binding of PU.1 and C/EBPβ at these locations tolerated changes in spacing ranging from 5-bp increase to >30-bp decrease. Collectively, these findings have implications for understanding mechanisms underlying enhancer selection and for the interpretation of non-coding genetic variation.

Data availability

All sequencing data generated during this study have been deposited in GEO under accession code GSE178080. For reviewer access, please go to https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE178080 and enter token inyjgyqcbrsnrwz into the box.

The following data sets were generated
The following previously published data sets were used

Article and author information

Author details

  1. Zeyang Shen

    Department of Cellular and Molecular Medicine, University of California, San Diego, La Jolla, United States
    Competing interests
    The authors declare that no competing interests exist.
  2. Rick Z Li

    Department of Cellular and Molecular Medicine, University of California, San Diego, La Jolla, United States
    Competing interests
    The authors declare that no competing interests exist.
  3. Thomas A Prohaska

    Department of Medicine, University of California, San Diego, La Jolla, United States
    Competing interests
    The authors declare that no competing interests exist.
  4. Marten A Hoeksema

    Department of Cellular and Molecular Medicine, University of California, San Diego, La Jolla, United States
    Competing interests
    The authors declare that no competing interests exist.
  5. Nathan J Spann

    Department of Cellular and Molecular Medicine, University of California, San Diego, La Jolla, United States
    Competing interests
    The authors declare that no competing interests exist.
  6. Jenhan Tao

    Department of Cellular and Molecular Medicine, University of California, San Diego, San Diego, United States
    Competing interests
    The authors declare that no competing interests exist.
  7. Gregory J Fonseca

    Department of Cellular and Molecular Medicine, University of California, San Diego, La Jolla, United States
    Competing interests
    The authors declare that no competing interests exist.
  8. Thomas Le

    Division of Biological Sciences, University of California, San Diego, La Jolla, United States
    Competing interests
    The authors declare that no competing interests exist.
  9. Lindsey K Stolze

    Department of Cellular and Molecular Medicine, University of Arizona, Tucson, United States
    Competing interests
    The authors declare that no competing interests exist.
  10. Mashito Sakai

    Department of Biochemistry and Molecular Biology, Nippon Medical School, Tokyo, Japan
    Competing interests
    The authors declare that no competing interests exist.
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0002-4908-2727
  11. Casey E Romanoski

    Department of Cellular and Molecular Medicine, University of Arizona, Tucson, United States
    Competing interests
    The authors declare that no competing interests exist.
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0002-0149-225X
  12. Christopher K Glass

    Department of Cellular and Molecular Medicine, University of California San Diego, La Jolla, United States
    For correspondence
    ckg@ucsd.edu
    Competing interests
    The authors declare that no competing interests exist.
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0003-4344-3592

Funding

National Institutes of Health (DK091183)

  • Christopher K Glass

National Institutes of Health (HL147835)

  • Christopher K Glass

Leducq Transatlantic Network (16CVD01)

  • Christopher K Glass

National Institutes of Health (T32DK007044)

  • Thomas A Prohaska

American Heart Association (postdoctoral grant)

  • Marten A Hoeksema

Netherlands Organization for Scientific Research (Rubicon grant)

  • Marten A Hoeksema

Amsterdam Cardiovascular Sciences Institute (postdoctoral grant)

  • Marten A Hoeksema

The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication.

Reviewing Editor

  1. Jessica K Tyler, Weill Cornell Medicine, United States

Ethics

Animal experimentation: Bone marrow cells were isolated from femurs and tibias of Cas9-expressing transgenic mice (Jackson Laboratory, No.028555) housed at the University of California San Diego animal facility on a 12-hour/12-hour light/dark cycle with free access to normal chow food and water. All of the mice were handled according to approved institutional animal care and use committee (IACUC) protocols (S01015) of the University of California San Diego to minimize pain and suffering.

Version history

  1. Preprint posted: April 3, 2020 (view preprint)
  2. Received: June 1, 2021
  3. Accepted: January 12, 2022
  4. Accepted Manuscript published: January 20, 2022 (version 1)
  5. Version of Record published: February 2, 2022 (version 2)

Copyright

© 2022, Shen et al.

This article is distributed under the terms of the Creative Commons Attribution License permitting unrestricted use and redistribution provided that the original author and source are credited.

Metrics

  • 2,124
    views
  • 240
    downloads
  • 3
    citations

Views, downloads and citations are aggregated across all versions of this paper published by eLife.

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Zeyang Shen
  2. Rick Z Li
  3. Thomas A Prohaska
  4. Marten A Hoeksema
  5. Nathan J Spann
  6. Jenhan Tao
  7. Gregory J Fonseca
  8. Thomas Le
  9. Lindsey K Stolze
  10. Mashito Sakai
  11. Casey E Romanoski
  12. Christopher K Glass
(2022)
Systematic analysis of naturally occurring insertions and deletions that alter transcription factor spacing identifies tolerant and sensitive transcription factor pairs
eLife 11:e70878.
https://doi.org/10.7554/eLife.70878

Share this article

https://doi.org/10.7554/eLife.70878

Further reading

    1. Chromosomes and Gene Expression
    Allison Coté, Aoife O'Farrell ... Arjun Raj
    Research Article

    Splicing is the stepwise molecular process by which introns are removed from pre-mRNA and exons are joined together to form mature mRNA sequences. The ordering and spatial distribution of these steps remain controversial, with opposing models suggesting splicing occurs either during or after transcription. We used single-molecule RNA FISH, expansion microscopy, and live-cell imaging to reveal the spatiotemporal distribution of nascent transcripts in mammalian cells. At super-resolution levels, we found that pre-mRNA formed clouds around the transcription site. These clouds indicate the existence of a transcription-site-proximal zone through which RNA move more slowly than in the nucleoplasm. Full-length pre-mRNA undergo continuous splicing as they move through this zone following transcription, suggesting a model in which splicing can occur post-transcriptionally but still within the proximity of the transcription site, thus seeming co-transcriptional by most assays. These results may unify conflicting reports of co-transcriptional versus post-transcriptional splicing.

    1. Chromosomes and Gene Expression
    2. Genetics and Genomics
    Maria L Adelus, Jiacheng Ding ... Casey E Romanoski
    Research Article

    Heterogeneity in endothelial cell (EC) sub-phenotypes is becoming increasingly appreciated in atherosclerosis progression. Still, studies quantifying EC heterogeneity across whole transcriptomes and epigenomes in both in vitro and in vivo models are lacking. Multiomic profiling concurrently measuring transcriptomes and accessible chromatin in the same single cells was performed on six distinct primary cultures of human aortic ECs (HAECs) exposed to activating environments characteristic of the atherosclerotic microenvironment in vitro. Meta-analysis of single-cell transcriptomes across 17 human ex vivo arterial specimens was performed and two computational approaches quantitatively evaluated the similarity in molecular profiles between heterogeneous in vitro and ex vivo cell profiles. HAEC cultures were reproducibly populated by four major clusters with distinct pathway enrichment profiles and modest heterogeneous responses: EC1-angiogenic, EC2-proliferative, EC3-activated/mesenchymal-like, and EC4-mesenchymal. Quantitative comparisons between in vitro and ex vivo transcriptomes confirmed EC1 and EC2 as most canonically EC-like, and EC4 as most mesenchymal with minimal effects elicited by siERG and IL1B. Lastly, accessible chromatin regions unique to EC2 and EC4 were most enriched for coronary artery disease (CAD)-associated single-nucleotide polymorphisms from Genome Wide Association Studies (GWAS), suggesting that these cell phenotypes harbor CAD-modulating mechanisms. Primary EC cultures contain markedly heterogeneous cell subtypes defined by their molecular profiles. Surprisingly, the perturbations used here only modestly shifted cells between subpopulations, suggesting relatively stable molecular phenotypes in culture. Identifying consistently heterogeneous EC subpopulations between in vitro and ex vivo models should pave the way for improving in vitro systems while enabling the mechanisms governing heterogeneous cell state decisions.