Spatiotemporal patterns of genetic diversity in the world’s coral reefs

  1. Remote Sensing Laboratories, Department of Geography, University of Zurich, Zurich, Switzerland
  2. Department of Chemistry, University of Zurich, Zurich, Switzerland

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Detlef Weigel
    Max Planck Institute for Biology Tübingen, Tübingen, Germany
  • Senior Editor
    Detlef Weigel
    Max Planck Institute for Biology Tübingen, Tübingen, Germany

Reviewer #1 (Public review):

Summary:

Despite setting global conservation goals for the conservation of genetic diversity, gathering enough genetic data to assess its current status or monitor change over time for all species would require substantial time and resources. Finding reliable proxies or predictors of genetic diversity change would be a valuable alternative. Selmoni and Schuman test for patterns and trends in coral reef animals' genetic diversity across space and time and evaluate how well environmental variables predict their genetic diversity.

Strengths:

The authors have compiled a large genomic dataset from 19 studies, 18 species, and 173 reefs, and use a standardized k-mer-based pipeline to process all datasets. The authors propose that using k-mers may be especially suitable for macrogenomic analyses because they are less computationally intense and do not require a reference genome. This is interesting because there are currently few examples of macrogenetics studies using genomic data, likely due in part to the difficulty associated with processing multiple disparate datasets in an efficient and standardized way.

Weaknesses:

I have concerns that the models don't fit the data structure, the data are not suitable to test for temporal change, and that the values being predicted do not reflect meaningful genetic diversity. The manuscript would also benefit from a more cohesive narrative structure, clearly stated research goals, and better engagement with previous literature on this topic.

Effects of time: The data are not suitable to test for temporal changes across all species. 11 of 19 datasets were only sampled over 1-2 years, and only 5 were sampled across 5 or more years. I don't know generation times for these species offhand, but this is generally too short to draw any conclusions about genetic diversity change over time. While it makes sense to include year in the models as a control variable, in my opinion these patterns should not be interpreted and removed from the discussion, or at least the conclusions should be tempered (for example the section "Are coral reef populations rapidly losing their genetic diversity?").

Environmental effects models: I'm not convinced that the geographic random intercept term (response variable) represents a local genetic distance that should be related to environments. How many species were sampled at each reef? If only 1 species is sampled, then the reef's average diversity reflects that species, not necessarily environmental effects. If multiple species are sampled at a reef, then I suspect this value is the average diversity of all the species sampled there. Knowing that species have different levels of diversity (e.g., Toczydlowski et al 2025 cited here), what does this metric mean and how is it affected by species composition vs environments? I do not see the value in predicting this variable without considering species identity.

Introduction: The macrogenetics literature is not well described in the introduction. Most of the cited studies indeed use raw data rather than summary statistics, and most do not aim to predict genetic diversity.

I wouldn't say that explaining more than 20% of variation in data is a major challenge in macrogenetics. The choice of how many, and which, predictors to use depends on the research question (e.g., on the spectrum of understanding to predicting, see Shmueli 2010 "To Explain or To Predict?"; Arif & MacNeil 2022 "Predictive models aren't for causal inference"). It's proposed here that explained variance can be increased by adding more, or more relevant, environmental predictors-but particularly for macrogenetics studies reusing data from different sources with different study and sampling designs and different species, much of the variation in the data will be due to these factors. That is why conditional R2s are typically quite high (>80%) in mixed models with species/study as random effects, see e.g. Clark & Pinsky 2024 or Karachaliou 2025. Having strong effects of particular variables would also require that all species respond to the predictor in the same way, which is not necessarily expected. Low explanatory capacity alone is not an issue if the goal is not prediction.

Reviewer #2 (Public review):

Summary:

The authors measured whether a phylogenetically wide sample of marine species was experiencing declining genetic diversity, as one might expect from widespread habitat threat, over a recent 2-decade time span. They next identified key environmental variables that are impacting genetic diversity within and between reefs. A key insight was to apply k-mer-based genetic distances to massively speed up the reanalysis of genetic data into a common pipeline, which is otherwise onerous. The manuscript ends by highlighting key seascape variables that were associated with increases or decreases with genetic diversity through time. The authors achieved their overall aims, though I remain unsure of how well the identified seascape variable-genetic diversity predictions can be generalized, as implied in the abstract.

Strengths:

A key strength of the study was its rigor in vetting the k-mer based distance metrics, checking whether they give population structure patterns (Figure S2) and correlated with nucleotide diversity. This surpasses previous studies that used a similar method. The modeling procedure of genetic diversity predictions from environmental variables is also rigorous and presented with some appropriate nuance.

Weaknesses:

I noted five weaknesses, listed below in order of potentially more severe at the top to more minor at the bottom.

First, only one k-mer distance metric was tested. There are many k-mer distance metrics that will potentially give different weight to different frequencies of polymorphism, like how Watterson's theta and nucleotide diversity weight polymorphisms, depending on their frequency. It would be interesting to see if the seascape variable predictions hold with a Jaccard or cosine distance metric, or if the observed results are purely restricted to the choice of Bray-Curtis distance. Another option would be to use mash, skmer, or (very recent development) re-skmer distances, which attempt to more directly approximate the average nucleotide identity between two sequence sets based on the k-mer sets of their reads while also being faster than traditional alignment. This would potentially give cleaner trends, seeing as the goal is to have a proxy for nucleotide diversity, with the downside of not including the impacts of non-SNP variation.

Second, it is unclear to me whether the number of species analyzed, 18, can accurately identify important environmental variables in early warning systems, as claimed in the abstract. While this likely represents the best available balance of evidence, it is worth highlighting the manuscript's note that the environment-diversity predictions did not scale across marine realms. This is likely a limitation of data availability, rather than a study design flaw, but is nonetheless important for readers to keep in mind. Would recommend that the abstract acknowledge this limitation.

Third, it is unclear to me how the included datasets compare in terms of genome-wide coverage. Figure S1 gives sequencing depths in terms of read number, but what is the range normalized for genome size - are the datasets 1x, 5x, 10x, on average, etc? This is probably most key to how the k-mer distance metrics will perform, because low coverage will make two samples appear artificially distant due to rarity of sampling the same k-mer multiple times. However, the correlation between k-mer distance and regular alignment-based SNP distance (Figure 2C) gives some confidence that this effect could be small.

Fourth, though k-mer distances may in some sense better capture the breadth of DNA sequence diversity, they lack a concrete interpretation of what loci may/may not be under selection as marine environments are increasingly threatened over time. This is a different question than what the manuscript tries to address, but is of interest to the field and is something that would be seemingly difficult to do with k-mers.

Finally, I had a more minor concern: the k-mer-based distances use only k-mers that are mapped to a reference genome. This is done to thoroughly remove contaminant sequences, which is important, but is a double-edged sword because there is potentially additional pangenomic variation that is real but does not map to a single reference. The authors state in the first paragraph of the discussion that this choice did not affect the results, but I did not find an associated analysis in the supplement or main text. It would be interesting to compare distance metrics based on screening against all known microbial+human genomes vs screening against the reference genomes.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation