Modeling workflow.

(A) Individual-level genomic data were compiled globally for various coral reef species, with known sampling site coordinates and years for each individual. Each sampling site was assigned to an ocean, ecoregion, sampling region (grouping sites within 100 km), and sampling reef (within 10 km). Bray-Curtis genetic distances (BCD) were then calculated between every pair of conspecific individuals. (B) A Generalized Linear Mixed Model (GLMM) assessed BCD variation across datasets, geographic regions, and sampling years. (C) The GLMM local effects on BCD were extracted for oceans, ecoregions, sampling regions, and sampling reefs. (D) For each sampling site, cumulative local effects on BCD (summing ocean, ecoregion, sampling region, and sampling reef effects) were calculated, alongside dozens of environmental variables characterizing seascape conditions. (E) A penalized regression model identified key seascape predictors of local BCD variation and generated spatial predictions for reefs worldwide.

Genomic datasets.

Overview of the 19 genomic datasets included in this study. For each dataset (row), the following information is provided: bibliographic reference (Ref.), species name, species type, ocean, ecoregion, number of samples before and after quality control (see Methods section on “Counting k-mer composition by sample”), number of sampling reefs (defined as sampling sites grouped within a 5 km radius), geographic range of the sampling reef distribution, and sampling period. For the total number of sampling sites, the number in parentheses indicates the number of unique sampling sites across all datasets.

Genomic distances across the world’s reefs.

(A) Geographic distribution of sampling locations across the 19 genomic datasets included in the study. Points represent sampling sites (N=173), linked by lines for each dataset (labeled from to 19, read below). The legend indicates the main taxonomic groups, with the number of datasets, sampling reefs, and ndividual samples shown in parentheses. For every dataset, the genetic distances between samples collected from different reefs (>10 km apart) and from the same reef (<10 km) are shown in (B). Genetic distances are calculated as Bray–Curtis distances (BCD) based on k-mer frequencies. (C) shows the correlation between within-reef BCD (calculated using k-mers) and pairwise nucleotide diversity (calculated from single nucleotide polymorphisms, SNPs). Correlation is reported using Spearman’s rank correlation coefficient (ρ). Datasets labels are: (1) Ancylomenes pedersoni from the Caribbean, (2) Stylophora pistillata and (3) Pocillopora verrucosa from the Red Sea, (4) Platygyra daedalea from the Persian Gulf, (5) Pocillopora damicornis (6) Pocillopora acuta and (7) Acropora millepora from New Caledonia, (8) Acropora palmata and (9) Bartholomea annulata from the Caribbean, (10) Pinctada fucata from the South China Sea, (11) Mobula alfredi from Hawaii, (12) Pomacanthus maculosus from the Indian Ocean, (13) Pterois volitans from the Caribbean, (14) Dascyllus trimaculatus from the Indian Ocean, (15) Siphamia tubifer from Japan, (16) Amphiprion bicinctus from the Red Sea, (17) Epinephelus striatus from Bahamas, (18 and 19) Carcharhinus amblyrhynchos from the Indo-Pacific.

Taxonomic and spatiotemporal variation in coral reef genetic distances.

(A) Stepwise model selection of variables explaining genetic distances (i.e. Bray-Curtis distances measured across 255,247 pairs of samples, from 18 coral reef species). Top: improvement in model fit (reduction of Akaike Inference Criterion, ΔAIC) relative to a null model including only dataset variation, for models incorporating taxonomy, sampling strategy, spatial autocorrelation, temporal autocorrelation, and latitude. Bottom names of variables included in each model. For the model with the best fit (lowest AIC, right side of A), (B-E) show the model-adjusted Bray–Curtis distances (BCDadj) for different variables, representing the marginal effect on genetic distances by dataset (B, points, with bars indicating the 95% confidence interval), geographic distance between sampling sites (C), temporal distance between years of sampling (D), and midpoint year of sampling (E). (F) BCDadj by the midpoint year of sampling for sample pairs collected from different geographic distances (<1km, 1-10km, 10-100km, >100km). In (C-E), points represent the model’s partial residuals.

Prediction of genetic distances across the world’s coral reefs.

(A) Cumulative local effects of geographic position o Bray–Curtis genetic distances at each sampling reef (N 136). Red indicates negative local effects o genetic distances; blue indicates positive effects. (B) Scaled effect sizes (with 95% confidence intervals) of the eight environmental variables retained in a penalized logistic model predicting local effects on genetic distances (in brackets: the spatial scale at which predictors were measured). Color reflects the direction of the effect (red negative, blue positive). Partia coefficients of determination (pR2) are shown on the right. (C) Predictive performance of the penalized logistic model under leave-one-region-out cross-validation, quantified using the Area Under the Curve (AUC). The black line shows the model based on environmental predictors; the grey line shows the performance of an analogous model that replaced environmental predictors with spatial predictors (distance-based Moran Eigenvector Maps, dbMEMs). (D) Global predictions of local effects o genetic distances cross coral reefs worldwide, based n the penalized logistic model Predicted probabilities of a positive effect (PΔBCD > 0) re shown: red ndicates likely positive effects (PΔBCD 1), blue negative effects (PΔBCD 0), and pale yellow indicates no predicted directional effect (PΔBCD 0.5). Bottom: Circular barplots display the scaled mean (± standard deviation) of the eight predictors shown in (B) across six focal reef regions. Bars are colored blue for variables associated with positive effects on genetic distances, and red for variables associated with negative effects.