The genetic control of rapid genome content divergence in Arabidopsis thaliana
Figures
K-mer analyses of the A. thaliana reference genome.
(A) Median K-mer frequency in repetitive and non-repetitive compartments of the reference genome. (B) Relationship between scaled 12-mer-based copy number estimates (y-axis) and simulated copy number change across 1000 simulations. Each gray line denotes a single simulation iteration, and the blue line indicates the median values across all iterations. (C) Relationship between 12-mer-based copy number estimates and annotation-based copy number estimates for 272 transposon families in the TAIR10 reference genome. The blue line indicates the best-fit line from a linear regression model. (D) Relationship between 12-mer abundances in Illumina reads (44 X mean coverage) and corresponding abundances in the genome assembly. Each hexagonal bin is colored by the number of observations within the bin. The blue line indicates the best-fit line from a linear regression model.
Distribution of R2 values from linear models to predict simulated copy number variation using 12-mer abundance.
Each model compares the actual and 12-mer estimated copy numbers for a single 1 kb sequence selected from the reference genome, with simulated copy number changes ranging from –1 (deletion) to 10 copies.
Box plots of residuals from linear regression between 12-mer abundances in the TAIR10 reference genome and high-throughput sequencing reads.
P-values were calculated using Wilcoxon rank-sum tests comparing median residuals between groups. Outliers are not shown.
Correlations between 12-mer frequencies in sampled or simulated high-throughput sequencing reads and the TAIR10 reference genome.
The dotted red line denotes 1 X coverage, which was used as the threshold for coverage-based filtering in the pipeline. Sampled reads were generated by downsampling Col-0 sequencing reads from the 1001 Genomes Project to each target coverage. Simulated reads were generated from the TAIR10 reference genome.
Pipeline for generating genome content profiles (GCPs) and estimating sequence copy number from 12-mers.
The pipeline takes raw sequencing reads as input, which are trimmed for quality and adapter sequences using Trimmomatic. The trimmed reads are then error-corrected with SPAdes and mapped to the organellar genomes using BWA-MEM to remove organellar reads. The remaining unmapped reads, presumed to originate from the nuclear genome, are used for 12-mer counting with Jellyfish. After 12-mers have been counted for all samples, counts are normalized first for GC bias and then for sequencing coverage, followed by linear model correction to remove sequencing center effects. The resulting GCPs are then used to estimate the copy number of focal sequences by taking the median normalized 12-mer count across all overlapping 12-mers comprising each focal sequence.
Quality control filters for the 12-mer counting pipeline.
Distributions of (A) the proportion of reads mapped to the TAIR10 genome per sequencing run, (B) sequencing coverage estimated from 12-mer counts, (C) GC content per sequencing run, (D) pairwise genetic distance (allele count) and (E) Spearman’s correlation between GC content and 12-mer abundance. Vertical red lines indicate hard filtering thresholds. In panels A, B, and D, samples with values below the threshold were filtered. In panels C and E, samples with values outside the ranges denoted by the vertical lines were filtered.
Effect of GC normalization on the GC content dependence of 12-mer profiles.
Tukey boxplots of 12-mer abundances binned by 12-mer GC content before (A) and after (B) GC normalization.
Effect of frequency normalization on 12-mer abundance distributions.
Empirical cumulative distribution functions of 12-mer abundances for 10,000 random K-mers before (A) and after (B) quantile-quantile (QQ) normalization.
Effect of sequencing center correction on principal component analysis of 12-mer profiles.
(A) Principal component analysis of 12-mer profiles before correction for sequencing center, with points colored by sequencing center. (B) Principal component analysis of 12-mer profiles after correction for sequencing center.
Patterns of copy number variability across the genome.
(A) Sequence abundance variability in non-overlapping 100 kb windows across the genome in relation to repeat density and major cytological features. sR denotes the standardized range (range/median) of 12-mer abundance across all individuals. (B) Number of individuals with putative copy number changes per repetitive sequence. Z-scores were calculated from 12-mer-derived copy number estimates, and y-axis values represent the number of z-score values greater than 3 (copy number increase, positive values) or less than –3 (copy number decrease, negative values) across sequences. Sequences colored teal are unclassified repeats.
Relationship between sR and reference genome sequence composition in 100 kb windows.
Abundance variability (log-standardized range; sR) per 100 kb window as a function of the proportion of (A) genic sequence and (B) repetitive sequence. Blue lines indicate linear model fits.
Abundance variability of sequences by class.
Tukey boxplots showing sequence abundance variability measured as the standardized range (sR). P-values were calculated from Wilcoxon–Mann–Whitney U tests.
Abundance outliers for retrotransposons by superfamily.
Each bar represents a retrotransposon family, and bar height indicates the number of accessions showing increased (positive values) or decreased (negative values) copy number relative to the population mean. Colors indicate superfamily.
Abundance outliers for DNA transposons by superfamily.
Each bar represents a DNA transposon family, and bar height indicates the number of accessions showing increased (positive values) or decreased (negative values) copy number relative to the population mean. Colors indicate superfamily.
Abundance outliers for satellites by subclass.
Each bar represents a representative satellite sequence, and bar height indicates the number of accessions showing increased (positive values) or decreased (negative values) copy number relative to the population mean. Colors indicate satellite subclass.
Abundance outliers for simple repeats by subclass.
Each bar represents a representative simple repeat sequence, and bar height indicates the number of accessions showing increased (positive values) or decreased (negative values) copy number relative to the population mean. Colors indicate simple repeat subclass.
Repeat copy number differences between populations.
(A) Principal component 1 of repeat copy number estimates partitioned by subpopulation. (B) Number of sequences with putative copy number changes per individual, grouped by subpopulation. Z-scores were calculated from 12-mer-derived copy number estimates, and y-axis values correspond to the number of z-score greater than 3 (copy number increase, positive values) or less than –3 (copy number decreases, negative value) across individuals. Abbreviated subpopulations are as follows: N.S., North Sweden; S.S., South Sweden; Af., Africa; R., relict; W. Eur., Western Europe; C. Eur., Central Europe; Ch., China; Ger., Germany; I.B.C., Italy Balkan Caucasus.
Distributions of first principal component (PC1) loadings by repeat class.
Each point represents the PC1 loading of a repetitive sequence, with points jittered to show the distribution within each class.
Distributions of first principal component (PC1) loadings by repeat subclass.
Each point represents the PC1 loading of a repetitive sequence, with points jittered to show the distribution within each class. This is the same data shown in Figure S15 but grouped by superfamily or subclass.
Frequency of Z score outliers by sequence class for the top 25 most divergent genome content profiles (GCPs).
Each column represents an accession, with labels colored by admixture group. Bar height indicates the number of sequences showing increased (positive values) or decreased (negative values) copy number relative to the population mean. Bars are filled according to repetitive sequence class.
Genomic regions associated with repeat copy number variation.
(A) Frequency of significant SNPs per 100 kb window across all genome-wide association study (GWAS) (gray), and GWAS in retrotransposons (blue), DNA transposons (red), satellites (green), and simple repeats (yellow). The pink highlighted region indicates centromeric and pericentromeric regions. (B) Proportion of all SNPs (black), significant SNPs from all GWAS (gray), and significant SNPs across GWAS by sequence class in the pericentromere.
Sequences with significant genome-wide association study (GWAS) SNPs in the pericentromere and centromere are more likely to occur in centromeric regions.
For each GWAS with at least 10 hits, we calculated the proportion of sequence annotations located in the pericentromere and centromere, and compared distributions based on whether significant SNPs from the GWAS were enriched in these regions.
Genome-wide association mapping of repeat abundance.
(A) Meta-genome-wide association study (GWAS) of repeat abundance. The dotted line indicates the Bonferroni-corrected significance threshold (ɑ=0.05). Labels correspond to candidate genes at each locus listed in Supplementary file 1,table 6. (B) GWAS of first principal component (PC1) of repeat abundance. The dotted line indicates the Bonferroni-corrected significance threshold (ɑ=0.05). (C) Frequency of significant meta-GWAS tag SNPs across individual GWAS, separated by sequence class. Bar colors indicate the effect of the minor allele on sequence abundance (i.e. positive β values indicate that the minor allele is associated with increased sequence copy number).
12-mer (A) and 10-mer (B) genome-wide association study (GWAS) for ATCOPIA78_I.
Manhattan plots showing GWAS results for the internal sequence of ONSEN using 12-mer- or 10-mer-based genome content profiles (GCPs). The dotted red line indicates the Bonferroni-adjusted significance threshold (ɑ=0.05).
12-mer and 10-mer genome-wide association study (GWAS for ATHATN2).
Manhattan plots showing GWAS results for the representative sequence of ATHATN2 using 12-mer (A) or 10-mer-based (B) genome content profiles (GCPs). Meta-GWAS candidate SNPs are highlighted in orange and labeled with either the associated candidate gene or the top SNP. The dotted red line indicates the Bonferroni-adjusted significance threshold (ɑ=0.05).
12-mer and 10-mer genome-wide association study (GWAS) for VANDAL1.
Manhattan plots showing GWAS results for the representative sequence of VANDAL1 using 12-mer (A) or 10-mer-based (B) genome content profiles (GCPs). The dotted red line indicates the Bonferroni-adjusted significance threshold (ɑ=0.05).
Evolutionary dynamics of alleles associated with repeat copy number variation.
(A) Site frequency spectra of meta-genome-wide association study (GWAS) tag SNPs, meta-GWAS tag SNPs stratified by predominant copy number effect, all GWAS-associated SNPs identified in this study, and AraGWAS-associated SNPs, compared to nonsynonymous and synonymous SNPs from the 1001 Genomes dataset. (B) Relationship between the median number of sequences with copy number changes per subpopulation (normalized to the Africa group) and the ratio of mean derived allele frequencies between nonsynonymous and synonymous SNPs.
Additional files
-
Supplementary file 1
Supplementary Tables.
- https://cdn.elifesciences.org/articles/108238/elife-108238-supp1-v1.xlsx
-
MDAR checklist
- https://cdn.elifesciences.org/articles/108238/elife-108238-mdarchecklist1-v1.docx