Epigenetics: Making the most of methylation
DNA methylation is a key mechanism used by higher eukaryotes to regulate gene expression. The addition of a methyl group to carbon atom number 5 within cytosine bases in DNA is known to repress the transcription of genes into messenger RNA molecules, thus reducing the production of the proteins coded by these genes. Most methylation occurs at CpG dinucleotides—cytosines that are paired with guanines—and these often cluster together to form CpG islands in the promoter regions of genes. In the late 1990s, it was discovered that transcription was repressed when methyl CpG binding proteins were recruited to methylated CpG islands (Hendrich and Bird, 1998).
Subsequent studies have confirmed that the binding of these proteins throughout the genome is proportional to the density of DNA methylation (Baubec et al., 2013), and have identified additional proteins with a high affinity for methylated CpG sites (reviewed in Defossez and Stancheva, 2011). Moreover, in recent years, other screening approaches based mainly on mass spectrometry have revealed that more proteins bind to methylated DNA than previously thought (Mittler et al., 2009; Bartke et al., 2010; Bartels et al., 2011; Spruijt et al., 2013). Now, in eLife, Heng Zhu and co-workers at the Johns Hopkins University School of Medicine—including Shaohui Hu as first author—use a high-throughput screening method to show that many human transcription factors also interact with genomic DNA sequences containing methylated CpG sites (Hu et al., 2013).
To this end, the Johns Hopkins researchers made use of a published protein microarray consisting of 1,321 transcription factors and 210 co-factors (Hu et al., 2009). Hu et al. incubated the array with 154 distinct human promoter sequences, each of which contained at least one methylated CpG dinucleotide.Their results revealed that 150 (97%) of the 154 methylated human promoter sequences showed specific binding to at least one protein on the microarray. Moreover, of the 1531 proteins, 47 (3%) showed binding to methylated cytosines within the promoters. Most of the proteins bound to methylated DNA in a sequence-dependent manner; however, a minority bound to many different methylated DNA probes, indicating that binding can sometimes occur independent of DNA sequence (Figure 1).
A number of transcription factors, including KLF4—a recently identified methyl-CpG binding protein (Spruijt et al., 2013)—interacted with methylated sequences that did not resemble their known consensus DNA binding motifs. Using a technique based on electrophoresis, Hu et al. showed that KLF4 binds methylated and non-methylated DNA in a non-competitive manner: this suggests that different domains of the protein may be responsible for each type of binding, which they duly confirmed using molecular modeling and mutagenesis studies.
The Johns Hopkins researchers then mined published ChIP-sequencing data from stem cells to identify the target DNA sequences of KLF4, and compared these with data on genome-wide DNA methylation. Strikingly, KLF4 binding appears to be bimodal in nature throughout the genome, with 38% of KLF4 binding sites showing less than 20% methylation, and 48% showing methylation levels over 80%. Finally, Hu et al. used ChIP-bisulfite sequencing, which makes it possible to determine the methylation status of each cytosine within a target DNA sequence, to confirm that KLF4 also binds to both methylated and non-methylated DNA in vivo.
Hu et al. only profiled a small fraction of the complete human methylome for interactions with transcription factors; further proteins capable of binding genomic methyl CpG sequences surely await identification. The same holds true for interactions with methylated non-CpG sequences such as methyl-CpA (cytosine adjacent to adenine), which are fairly abundant in embryonic stem cells (Ramsahoye et al., 2000; Lister et al., 2009). To determine the physiological relevance of these interactions, it will be important to deduce the affinity with which proteins bind these sequences compared to their known targets; initial experiments along these lines are presented in the current eLife paper. Furthermore, recent evidence suggests that non-methylated CpG islands recruit activator proteins, many of which contain a CXXC motif (reviewed in Long et al., 2013). The transcription factor microarray approach used by the Johns Hopkins team, combined with quantitative mass spectrometry-based technology (Spruijt et al., 2013), could thus be used to identify the complete cellular complement of proteins that bind specifically to non-methylated CpG islands.
Finally, this study and other recently published papers force us to reconsider the mechanism(s) via which CpG methylation regulates transcription. Although DNA methylation is generally considered to be a repressive epigenetic modification, experiments presented by Hu et al. suggest that in some cases, methylation of a given promoter sequence can result in activation of transcription. Moreover, other work has revealed a temporal uncoupling of DNA methylation and transcriptional repression during Xenopus embryogenesis (Bogdanovic et al., 2011). Further experiments are therefore required to determine whether the functional readout of CpG methylation is affected by the repertoire and abundance of different DNA methylation ‘readers’ acting at any given time in a cell or a developing organism.
References
-
Biological functions of methyl-CpG-binding proteinsProg Mol Biol Transi Sci 101:377–398.https://doi.org/10.1016/B978-0-12-387685-0.00012-3
-
Identification and characterization of a family of mammalian methyl-CpG binding proteinsMol Cell Biol 18:6538–6547.
-
ZF-CxxC domain-containing proteins, CpG islands and the chromatin connectionBiochem Soc trans 41:727–740.https://doi.org/10.1042/BST20130028
Article and author information
Author details
Publication history
Copyright
© 2013, Vermeulen
This article is distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use and redistribution provided that the original author and source are credited.
Metrics
-
- 523
- views
-
- 57
- downloads
-
- 2
- citations
Views, downloads and citations are aggregated across all versions of this paper published by eLife.
Download links
Downloads (link to download the article as PDF)
Open citations (links to open the citations from this article in various online reference manager services)
Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)
Further reading
-
- Biochemistry and Chemical Biology
The conformational ensemble and function of intrinsically disordered proteins (IDPs) are sensitive to their solution environment. The inherent malleability of disordered proteins, combined with the exposure of their residues, accounts for this sensitivity. One context in which IDPs play important roles that are concomitant with massive changes to the intracellular environment is during desiccation (extreme drying). The ability of organisms to survive desiccation has long been linked to the accumulation of high levels of cosolutes such as trehalose or sucrose as well as the enrichment of IDPs, such as late embryogenesis abundant (LEA) proteins or cytoplasmic abundant heat-soluble (CAHS) proteins. Despite knowing that IDPs play important roles and are co-enriched alongside endogenous, species-specific cosolutes during desiccation, little is known mechanistically about how IDP-cosolute interactions influence desiccation tolerance. Here, we test the notion that the protective function of desiccation-related IDPs is enhanced through conformational changes induced by endogenous cosolutes. We find that desiccation-related IDPs derived from four different organisms spanning two LEA protein families and the CAHS protein family synergize best with endogenous cosolutes during drying to promote desiccation protection. Yet the structural parameters of protective IDPs do not correlate with synergy for either CAHS or LEA proteins. We further demonstrate that for CAHS, but not LEA proteins, synergy is related to self-assembly and the formation of a gel. Our results suggest that functional synergy between IDPs and endogenous cosolutes is a convergent desiccation protection strategy seen among different IDP families and organisms, yet the mechanisms underlying this synergy differ between IDP families.
-
- Biochemistry and Chemical Biology
- Stem Cells and Regenerative Medicine
Human induced pluripotent stem cells (hiPSCs) have great potential to be used as alternatives to embryonic stem cells (hESCs) in regenerative medicine and disease modelling. In this study, we characterise the proteomes of multiple hiPSC and hESC lines derived from independent donors and find that while they express a near-identical set of proteins, they show consistent quantitative differences in the abundance of a subset of proteins. hiPSCs have increased total protein content, while maintaining a comparable cell cycle profile to hESCs, with increased abundance of cytoplasmic and mitochondrial proteins required to sustain high growth rates, including nutrient transporters and metabolic proteins. Prominent changes detected in proteins involved in mitochondrial metabolism correlated with enhanced mitochondrial potential, shown using high-resolution respirometry. hiPSCs also produced higher levels of secreted proteins, including growth factors and proteins involved in the inhibition of the immune system. The data indicate that reprogramming of fibroblasts to hiPSCs produces important differences in cytoplasmic and mitochondrial proteins compared to hESCs, with consequences affecting growth and metabolism. This study improves our understanding of the molecular differences between hiPSCs and hESCs, with implications for potential risks and benefits for their use in future disease modelling and therapeutic applications.