Predicting geographic location from genetic variation with deep neural networks
Abstract
Most organisms are more closely related to nearby than distant members of their species, creating spatial autocorrelations in genetic data. This allows us to predict the location of a genetic sample by comparing it to a set of samples of known geographic origin. Here we describe a deep learning method, which we call Locator, to accomplish this task faster and more accurately than existing approaches. In simulations, Locator infers sample location to within 4.1 generations of dispersal and runs at least an order of magnitude faster than a recent model-based approach. We leverage Locator's computational efficiency to predict locations separately in windows across the genome, which allows us to both quantify uncertainty and describe the mosaic ancestry and patterns of geographic mixing that characterize many populations. Applied to whole-genome sequence data from Plasmodium parasites, Anopheles mosquitoes, and global human populations, this approach yields median test errors of 16.9km, 5.7km, and 85km, respectively.
Data availability
Locator is implemented as a command-line program written in Python: www.github.com/kern-lab/locator. SNP calls for the Anopheles dataset are available at https://www.malariagen.net/data/ag1000g-phase1-ar3, for P. falciparum at https://www.malariagen.net/resource/26,and for the HGDP at ftp://ngs.sanger.ac.uk/production/hgdp. Code to run continuous-space simulations can be found at https://github.com/kern-lab/spaceness/blob/master/slim_recipes/spaceness.slim. This publication uses data from the MalariaGEN Plasmodium falciparum Community Project as described in Pearson et al. (2019). Statistical analyses and many plots were produced in R (R Core Team, 2018).
-
Ag1000G phase 1 AR3 data releaseMalariaGEN, http://www.malariagen.net/data/ag1000g-phase1-AR3.
-
Plasmodium falciparum community project version 6 data releaseMalariaGEN, https://www.malariagen.net/resource/26.
-
Insights into human genetic variation and population history from 929 diverse genomesHGDP, ftp://ngs.sanger.ac.uk/production/hgdp.
Article and author information
Author details
Funding
National Institutes of Health (R01GM117241)
- CJ Battey
- Andrew D Kern
The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication.
Copyright
© 2020, Battey et al.
This article is distributed under the terms of the Creative Commons Attribution License permitting unrestricted use and redistribution provided that the original author and source are credited.
Metrics
-
- 9,551
- views
-
- 845
- downloads
-
- 80
- citations
Views, downloads and citations are aggregated across all versions of this paper published by eLife.
Download links
Downloads (link to download the article as PDF)
Open citations (links to open the citations from this article in various online reference manager services)
Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)
Further reading
-
- Ecology
- Evolutionary Biology
While host phenotypic manipulation by parasites is a widespread phenomenon, whether tumors, which can be likened to parasite entities, can also manipulate their hosts is not known. Theory predicts that this should nevertheless be the case, especially when tumors (neoplasms) are transmissible. We explored this hypothesis in a cnidarian Hydra model system, in which spontaneous tumors can occur in the lab, and lineages in which such neoplastic cells are vertically transmitted (through host budding) have been maintained for over 15 years. Remarkably, the hydras with long-term transmissible tumors show an unexpected increase in the number of their tentacles, allowing for the possibility that these neoplastic cells can manipulate the host. By experimentally transplanting healthy as well as neoplastic tissues derived from both recent and long-term transmissible tumors, we found that only the long-term transmissible tumors were able to trigger the growth of additional tentacles. Also, supernumerary tentacles, by permitting higher foraging efficiency for the host, were associated with an increased budding rate, thereby favoring the vertical transmission of tumors. To our knowledge, this is the first evidence that, like true parasites, transmissible tumors can evolve strategies to manipulate the phenotype of their host.
-
- Evolutionary Biology
- Microbiology and Infectious Disease
Accurate estimation of the effects of mutations on SARS-CoV-2 viral fitness can inform public-health responses such as vaccine development and predicting the impact of a new variant; it can also illuminate biological mechanisms including those underlying the emergence of variants of concern. Recently, Lan et al. reported a model of SARS-CoV-2 secondary structure and its underlying dimethyl sulfate reactivity data (Lan et al., 2022). I investigated whether base reactivities and secondary structure models derived from them can explain some variability in the frequency of observing different nucleotide substitutions across millions of patient sequences in the SARS-CoV-2 phylogenetic tree. Nucleotide basepairing was compared to the estimated ‘mutational fitness’ of substitutions, a measurement of the difference between a substitution’s observed and expected frequency that is correlated with other estimates of viral fitness (Bloom and Neher, 2023). This comparison revealed that secondary structure is often predictive of substitution frequency, with significant decreases in substitution frequencies at basepaired positions. Focusing on the mutational fitness of C→U, the most common type of substitution, I describe C→U substitutions at basepaired positions that characterize major SARS-CoV-2 variants; such mutations may have a greater impact on fitness than appreciated when considering substitution frequency alone.