Dynamically evolving novel overlapping gene as a factor in the SARS-CoV-2 pandemic
Abstract
Understanding the emergence of novel viruses requires an accurate and comprehensive annotation of their genomes. Overlapping genes (OLGs) are common in viruses and have been associated with pandemics, but are still widely overlooked. We identify and characterize ORF3d, a novel OLG in SARS-CoV-2 that is also present in Guangxi pangolin-CoVs but not other closely related pangolin-CoVs or bat-CoVs. We then document evidence of ORF3d translation, characterize its protein sequence, and conduct an evolutionary analysis at three levels: between taxa (21 members of Severe acute respiratory syndrome-related coronavirus), between human hosts (3978 SARS-CoV-2 consensus sequences), and within human hosts (401 deeply sequenced SARS-CoV-2 samples). ORF3d has been independently identified and shown to elicit a strong antibody response in COVID-19 patients. However, it has been misclassified as the unrelated gene ORF3b, leading to confusion. Our results liken ORF3d to other accessory genes in emerging viruses and highlight the importance of OLGs.
Data availability
All data generated or analyzed during this study are included in the manuscript and supplement. Scripts and source data for all analyses and figures are provided on GitHub at https://github.com/chasewnelson/SARS-CoV-2-ORF3d and Zenodo at https://zenodo.org/record/4052729.
-
Proteome and Translatome of SARS-CoV-2 infected cellsPRIDE database, PXD017710.
-
Vero cells infected with SARS CoV 2 no quantitation slices 1-10 of 20; vero cells infected with SARS CoV2 slices 11-20 of 20 slicesZenodo, 10.5281/zenodo.3722590, 10.5281/zenodo.3722596.
-
Proteomics of SARS-CoV and SARS-CoV-2 infected cellsPRIDE database, PXD018581.
-
Decoding SARS-CoV-2 coding capacityGEO database, sample IDs: SRR11713366, SRR11713367, SRR11713368, SRR11713369 from GSE149973.
-
Assays and Merits of Proteomics for SARS-CoV-2 Research and TestingPRIDE database, PXD019645.
Article and author information
Author details
Funding
Academia Sinica (Postdoctoral Research Fellowship)
- Chase W Nelson
National Philanthropic Trust (Grant)
- Zachary Ardern
University of Wisconsin-Madison (John D. MacArthur Professorship Chair)
- Tony L Goldberg
National Science Foundation (IOS grants #1755370 and #1758800)
- Sergios-Orestis Kolokotronis
The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication.
Reviewing Editor
- Antonis Rokas, Vanderbilt University, United States
Publication history
- Received: June 3, 2020
- Accepted: September 30, 2020
- Accepted Manuscript published: October 1, 2020 (version 1)
- Version of Record published: November 10, 2020 (version 2)
Copyright
© 2020, Nelson et al.
This article is distributed under the terms of the Creative Commons Attribution License permitting unrestricted use and redistribution provided that the original author and source are credited.
Metrics
-
- 15,894
- Page views
-
- 1,414
- Downloads
-
- 37
- Citations
Article citation count generated by polling the highest count across the following sources: Crossref, PubMed Central, Scopus.
Download links
Downloads (link to download the article as PDF)
Open citations (links to open the citations from this article in various online reference manager services)
Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)
Further reading
-
- Evolutionary Biology
- Genetics and Genomics
Despite decades of research, knowledge about the genes that are important for development and function of the mammalian eye and are involved in human eye disorders remains incomplete. During mammalian evolution, mammals that naturally exhibit poor vision or regressive eye phenotypes have independently lost many eye-related genes. This provides an opportunity to predict novel eye-related genes based on specific evolutionary gene loss signatures. Building on these observations, we performed a genome-wide screen across 49 mammals for functionally uncharacterized genes that are preferentially lost in species exhibiting lower visual acuity values. The screen uncovered several genes, including SERPINE3, a putative serine proteinase inhibitor. A detailed investigation of 381 additional mammals revealed that SERPINE3 is independently lost in 18 lineages that typically do not primarily rely on vision, predicting a vision-related function for this gene. To test this, we show that SERPINE3 has the highest expression in eyes of zebrafish and mouse. In the zebrafish retina, serpine3 is expressed in Müller glia cells, a cell type essential for survival and maintenance of the retina. A CRISPR-mediated knockout of serpine3 in zebrafish resulted in alterations in eye shape and defects in retinal layering. Furthermore, two human polymorphisms that are in linkage with SERPINE3 are associated with eye-related traits. Together, these results suggest that SERPINE3 has a role in vertebrate eyes. More generally, by integrating comparative genomics with experiments in model organisms, we show that screens for specific phenotype-associated gene signatures can predict functions of uncharacterized genes.
-
- Evolutionary Biology
- Microbiology and Infectious Disease
Gene duplication is crucial to generating novel signaling pathways during evolution. However, it remains unclear how the redundant proteins produced by gene duplication ultimately acquire new interaction specificities to establish insulated paralogous signaling pathways. Here, we used ancestral sequence reconstruction to resurrect and characterize a bacterial two-component signaling system that duplicated in α-proteobacteria. We determined the interaction specificities of the signaling proteins that existed before and immediately after this duplication event and then identified key mutations responsible for establishing specificity in the two systems. Just three mutations, in only two of the four interacting proteins, were sufficient to establish specificity of the extant systems. Some of these mutations weakened interactions between paralogous systems to limit crosstalk. However, others strengthened interactions within a system, indicating that the ancestral interaction, although functional, had the potential to be strengthened. Our work suggests that protein-protein interactions with such latent potential may be highly amenable to duplication and divergence.