K-mer analyses of the A. thaliana reference genome

(A) Median K-mer frequency in repetitive and non-repetitive compartments of the reference genome. (B) Relationship between scaled 12-mer based copy number estimates (y-axis) and simulated copy number change across 1,000 simulations. Each gray line denotes a single simulation iteration, and the blue line indicates the median values across all iterations. (C) Relationship between 12-mer-based copy number estimates and annotation-based copy number estimates for 272 transposon families in the TAIR10 reference genome. The blue line indicates the best-fit line from a linear regression model. (D) Relationship between 12-mer abundances in Illumina reads (44 × mean coverage) and corresponding abundances in the genome assembly. Each hexagonal bin is colored by the number of observations within the bin. The blue line indicates the best-fit line from a linear regression model.

Patterns of copy number variability across the genome

(A) Sequence abundance variability in non-overlapping 100 kb windows across the genome in relation to repeat density and major cytological features. sR denotes the standardized range (range / median) of 12-mer abundance across all individuals. (B) Number of individuals with putative copy number changes per repetitive sequence. Z-scores were calculated from the 12-mer derived copy number estimates, and y-axis values represent the number of z-score values greater than 3 (copy number increase, positive value) or less than -3 (copy number decrease, negative value) across sequences. Sequences colored teal are unclassified repeats.

Repeat copy number differences between populations

(A) Principal component 1 of repeat copy number estimates partitioned by subpopulation. (B) Number of sequences with putative copy number change per individual grouped by subpopulation. Z-scores were calculated from 12-mer derived copy number estimates and y-axis values correspond to the count of z-score values greater than 3 (copy number increase, positive value) or less than -3 (copy number decrease, negative value) across individuals. Abbreviated subpopulations are as follows: N.S. is North Sweden, S.S. is South Sweden, Af. is Africa, R. is relict, W. Eur. is Western Europe, C. Eur. is Central Europe, Ch. is China, Ger. is Germany, I.B.C. is Italy Balkan Caucasus.

Genomic regions associated with repeat copy number variation.

(A) Frequency of significant SNPs per 100 kb windows across all GWAS (gray), or GWAS in retrotransposons (blue), DNA transposons (red), satellites (green), and simple repeats (yellow), respectively. Pink highlighted region indicates centromeric and pericentromeric regions. (B) Proportion of all SNPs (black), significant SNPs from all GWAS (gray), and significant SNPs across GWAS by sequence class in the pericentromere.

Genome-wide association mapping of repeat abundance.

(A) Meta-GWAS of repeat abundance. The dotted line indicates the significance threshold at Bonferroni-corrected alpha = 0.05. Labels correspond to candidate genes at each locus listed in Table S6. (B) GWAS of PC 1 of repeat abundance. The dotted line indicates significance threshold at Bonferroni-corrected alpha 0.05. (C) Frequency of significant meta-GWAS tag SNPs across individual GWAS separated by sequence class. Bar colors indicate the effect of the minor allele on sequence abundance (i.e. positive beta corresponds to minor allele associated with increased sequence copy number).

Evolutionary dynamics of alleles associated with copy number change.

(A) Site frequency spectra of meta-GWAS tag SNPs, meta-GWAS tag SNPs stratified by predominant copy number effect, all GWAS-associated SNPs identified in this study, and AraGWAS-associated SNPs, compared to nonsynonymous and synonymous SNPs from the 1001 Genomes dataset. (B) Relationship between the median number of sequences with copy number change per subpopulation (normalized to the Africa group) and the ratio of mean derived allele frequencies between nonsynonymous and synonymous SNPs.