Spermatogenesis: A hotspot for new genes

Single-cell RNA-sequencing in fruit flies gives an unprecedented picture of how new genes are expressed during the formation of sperm.
  1. Anne-Marie Dion-Côté  Is a corresponding author
  1. Université de Moncton, Canada

New genes can be produced in a number of different ways. Existing genes can be duplicated, while frameshift mutations may change the way a cell reads a sequence, which could lead to new proteins. 'De novo' genes can also be created from previously non-coding sequences. It was thought until recently that the emergence of these genes was extremely rare, but the advent of modern genomics and transcriptomics has revealed that this is not the case (Schlötterer, 2015; Jacob, 1977). In multicellular organisms, only a small proportion of the genome codes for proteins, yet a large fraction of the rest of the genome is actively transcribed and translated (Clark et al., 2011; Ruiz-Orera et al., 2018). These non-coding sequences provide the raw material for natural selection to act upon and create new genes: considering that most of the genome is non-coding in multicellular organisms, the potential for de novo genes to emerge is enormous.

De novo genes tend to be expressed in the testis and involved in male reproductive processes (Levine et al., 2006; Begun et al., 2007; Zhao et al., 2014). However, many of these genes are not fixed within a species, which means they are susceptible to disappear. In addition, a large number of genes associated with male reproduction are rapidly evolving (Swanson et al., 2001). These observations are surprising given that sperm formation is a complex, highly conserved process. It is possible that de novo genes arise more often in the testis because future sperm cells have a permissive DNA state which would allow transcription to proceed less specifically than in other tissues, exposing non-coding sequences to selection (Kaessmann, 2010). However, this hypothesis is notoriously difficult to test because the testis is a highly heterogeneous tissue: for example, germ cells pass through many different stages before they mature into sperm.

Single-cell RNA sequencing is a powerful technology that circumvents the challenges associated with tissue heterogeneity by revealing the expression profile of individual cells. It can be harnessed to study mixed cell populations in tumors, or to track transcriptional dynamics during complex developmental processes such as sperm formation. This mechanism, also known as spermatogenesis, requires a pool of germ stem cells to undergo a series of mitotic and meiotic divisions to finally produce mature sperm (Figure 1). Now, in eLife, Li Zhao and colleagues at the Rockefeller University – including Evan Witt as first author, Sigi Benjamin and Nicolas Svetec – report having leveraged single-cell RNA sequencing to investigate how de novo genes are expressed in fly testis during the different stages of spermatogenesis (Witt et al., 2019).

New genes are expressed differently depending on spermatogenesis stages.

In the testis of fruit flies, the creation of mature sperm, or spermatogenesis, starts with a germline stem cell going through several rounds of mitosis to form early spermatocytes. After meiosis, these cells become late spermatocytes, which then develop into early and late spermatids. Witt et al. show that fixed de novo genes, which emerge from non-coding sequences, are expressed during mid-spermatogenesis, in particular in early spermatocytes. In contrast, other types of new genes, for example which come from gene duplication, are expressed at different stages.

Image credit: Witt et al., 2019; adapted from Figure 1A (CC BY 4.0).

First, the Rockefeller team was able to categorize the cells as germline stem cells, somatic cells, or cells going through specific stages of maturation by analyzing the expression of well-known spermatogenesis marker genes. In addition, the researchers used overall transcriptional changes to reconstruct an inferred ‘pseudotime’, a roadmap of the different stages that cells go through as they develop into sperm. Together, these approaches allowed Witt et al. to thoroughly document the gene expression profiles of specific cell types during spermatogenesis.

In particular, they found that fixed de novo genes (that is, de novo genes that are unlikely to disappear from a species) were expressed differently depending on developmental stage. For example, they were more often switched on in meiotic germ cells, a pattern reminiscent of canonical genes expressed only in testes. However, unfixed de novo genes were less expressed in germline stem cells, and enriched in early sperm cells (which have gone through meiosis). These results indicate that de novo genes could be expressed during meiosis, and that the expression pattern of a de novo gene may impact its probability to reach fixation.

Next, Witt et al. compared the expression profiles of de novo and recent duplicate genes. The expression pattern of almost half of the derived duplicate genes was biased towards early and late spermatogenesis, a pattern never observed for parental genes. On the contrary, the expression of the majority of de novo genes was biased towards mid-spermatogenesis. These observations suggest that de novo and young duplicate genes are regulated differently, and that the mode of emergence of a new gene constrains how it is expressed.

Finally, Witt et al. developed a method to infer the presence of mutations from single-cell RNA sequencing data. Combined with the cell type and pseudotime information, this allowed them to estimate the mutational load – the amount of potentially deleterious mutations – over spermatogenesis. They found that the mutational load tends to decrease over pseudotime, suggesting that DNA damage was repaired or that carrier cells were eliminated. While this finding requires further validation, it opens the door to many questions related to DNA repair dynamics, selection pressure within the testis and even male fertility.

As evolution largely relies on the mutation rate in the germline, the work by Witt et al. elegantly highlights how single-cell sequencing technologies can start to address long-standing questions in evolutionary biology. One of the upcoming challenges is now to integrate these highly multidimensional data with the complex molecular pathways involved in reproduction and development.


Article and author information

Author details

  1. Anne-Marie Dion-Côté

    Anne-Marie Dion-Côté is in the Département de Biologie, Université de Moncton, Moncton, Canada

    For correspondence
    Competing interests
    No competing interests declared
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0002-8656-4127

Publication history

  1. Version of Record published: August 29, 2019 (version 1)


© 2019, Dion-Côté

This article is distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use and redistribution provided that the original author and source are credited.


  • 2,350
    Page views
  • 226
  • 3

Article citation count generated by polling the highest count across the following sources: Crossref, PubMed Central, Scopus.

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Anne-Marie Dion-Côté
Spermatogenesis: A hotspot for new genes
eLife 8:e50136.

Further reading

    1. Evolutionary Biology
    Priscila S Rothier, Anne-Claire Fabre ... Anthony Herrel
    Research Article

    Vertebrate limb morphology often reflects the environment due to variation in locomotor requirements. However, proximal and distal limb segments may evolve differently from one another, reflecting an anatomical gradient of functional specialization that has been suggested to be impacted by the timing of development. Here we explore whether the temporal sequence of bone condensation predicts variation in the capacity of evolution to generate morphological diversity in proximal and distal forelimb segments across more than 600 species of mammals. Distal elements not only exhibit greater shape diversity, but also show stronger within-element integration and, on average, faster evolutionary responses than intermediate and upper limb segments. Results are consistent with the hypothesis that late developing distal bones display greater morphological variation than more proximal limb elements. However, the higher integration observed within the autopod deviates from such developmental predictions, suggesting that functional specialization plays an important role in driving within-element covariation. Proximal and distal limb segments also show different macroevolutionary patterns, albeit not showing a perfect proximo-distal gradient. The high disparity of the mammalian autopod, reported here, is consistent with the higher potential of development to generate variation in more distal limb structures, as well as functional specialization of the distal elements.

    1. Developmental Biology
    2. Evolutionary Biology
    James W Truman, Jacquelyn Price ... Tzumin Lee
    Research Article

    We have focused on the mushroom bodies (MB) of Drosophila to determine how the larval circuits are formed and then transformed into those of the adult at metamorphosis. The adult MB has a core of thousands of Kenyon neurons; axons of the early-born g class form a medial lobe and those from later-born a'b' and ab classes form both medial and vertical lobes. The larva, however, hatches with only g neurons and forms a vertical lobe 'facsimile' using larval-specific axon branches from its g neurons. Computations by the MB involves MB input (MBINs) and output (MBONs) neurons that divide the lobes into discrete compartments. The larva has 10 such compartments while the adult MB has 16. We determined the fates of 28 of the 32 types of MBONs and MBINs that define the 10 larval compartments. Seven larval compartments are eventually incorporated into the adult MB; four of their larval MBINs die, while 12 MBINs/MBONs continue into the adult MB although with some compartment shifting. The remaining three larval compartments are larval specific, and their MBIN/MBONs trans-differentiate at metamorphosis, leaving the MB and joining other adult brain circuits. With the loss of the larval vertical lobe facsimile, the adult vertical lobes, are made de novo at metamorphosis, and their MBONs/MBINs are recruited from the pool of adult-specific cells. The combination of cell death, compartment shifting, trans-differentiation, and recruitment of new neurons result in no larval MBIN-MBON connections persisting through metamorphosis. At this simple level, then, we find no anatomical substrate for a memory trace persisting from larva to adult. For the neurons that trans-differentiate, our data suggest that their adult phenotypes are in line with their evolutionarily ancestral roles while their larval phenotypes are derived adaptations for the larval stage. These cells arise primarily within lineages that also produce permanent MBINs and MBONs, suggesting that larval specifying factors may allow information related to birth-order or sibling identity to be interpreted in a modified manner in these neurons to cause them to adopt a modified, larval phenotype. The loss of such factors at metamorphosis, though, would then allow these cells to adopt their ancestral phenotype in the adult system.