Figures and data

Overview of datasets, preprocessing pipeline, model architecture, and performance metrics for the fungal language model (Shorkie LM).
(A) Schematic of the Shorkie LM architecture. (B) Four datasets employed: single S. cerevisiae genome (R64_yeast, species-level), 80 S. cerevisiae strains (80_strains, strain-level), 165 Saccharomycetales genomes (165_Saccharomycetales, order-level) with Saccharomycetales highlighted in light green, and 1,341 fungal genomes spanning the kingdom (1341_Fungal, kingdom-level), including common mushrooms such as oyster mushroom (Pleurotus ostreatus), shiitake (Lentinula edodes), white button mushroom (Agaricus bisporus), and black truffle (Tuber melanosporum), as well as fission yeast (Schizosaccharomyces pombe), and brewer’s yeast (Saccharomyces cerevisiae). (C) Representative genome distance dot plots for selected genomes from each dataset, with the x-axis representing the R64 S. cerevisiae genome and the y-axis representing the comparison genome. (D) Mash distance between R64 S. cerevisiae genome and genomes in 80_strains and 165_Saccharomycetales. (E) Data preprocessing pipeline converting raw genomic data into tensors and labels for Shorkie LM training, validation, and testing. (F) Validation loss progression across training steps. (G) Comparison of test set perplexity in genic and intergenic regions across four model variants.

Shorkie LM identifies conserved transcription factor binding motifs across fungal genomes.
(A) Position probability matrix (PPM) reconstruction from DNA sequences using the fungal language model Shorkie LM. (B) Comparative analysis of SMT3 promoter predictions (chrIV:1,469,090-1,469,198) by Shorkie LM and the Speciesaware DNA LM22,52, highlighting key motifs including poly(dA:dT), Cbf1, Tye7, and Reb1. (C) Summary of known motifs detected by Shorkie LM across six datasets: (1) reference S. cerevisiae genome; (2) four randomly selected S. cerevisiae strains; (3) five genomes from the Saccharomycetales order; (4) four genomes from the Ascomycota phylum; (5) four genomes from the Orbiliales order; and (6) four genomes from the Schizosaccharomycetales order. TF-MoDISco-identified motifs include TATA-binding protein, 5′ splice site (donor), branch point, Cbf1p, Reb1.1, Snf1.1, Mcm1.1, Rap1.1, Sfp1.2, Abf1.1 and Dot6. (D) Histograms depicting enrichment of TF-MoDISco-identified motifs upstream of transcription start sites (TSS) relative to background distributions in S. cerevisiae, and enrichment of 5’ splice sites (donors) and branch points within genic regions. (E) t-SNE embeddings of different genomic elements from the first self-attention layer of Shorkie LM.

Shorkie architecture and RNA-seq prediction performance across multiple scales.
(A) Shorkie architecture: U-Net model with eight transformer blocks. All layers inherit pretrained Shorkie LM weights; task-specific output heads (blue) predict perturbation timepoint RNA-seq (n = 3,053), 1000-strain RNA-seq (n = 1,014), ChIP-exo (n = 1,128) and ChIP-MNase histone marks (n = 20). (B) Yeast cells were grown to steady state, as determined by culture density, prior to addition of b-estradiol to the culture and subsequent sampling. (C) Distribution of bin-level Pearson’s R on held-out test data for each track type, comparing Shorkie and Shorkie_Random_Init. (D-G) Scatter plots comparing Shorkie and Shorkie_Random_Init for RNA-seq tracks at (D) bin-level Pearson’s R; (E) gene-level Pearson’s R; (F) quantile-normalized and mean-centered gene-level Pearson’s R; (G) gene-by-gene, track-level Pearson’s R. (H-J) RNA-seq coverage snapshots of S. cerevisiae test set gene loci: (H) chrVII:362,180–366,023 (RPL7A); (I) chrIV:305,657–310,505 (RPS16B and RPL13A); and (J) chrVII:495,374–499,965 (EFM5).

Shorkie uses promoter and splicing motifs learned during pretraining.
(A-C) Promoter regions (−450 to +50 bp relative to the TSS; 500 bp total) of RPL26A (chrXII:818,862–819,362), FUN12 (chrI:75,977–76,477), and KRE33 (chrXIV:374,871–375,371). Rows 1-3 show DNA logos from Shorkie LM, and ISM maps from Shorkie (fine-tuned) and Shorkie_Random_Init (no self-supervision pretraining). Row 4 shows gene annotations from IGV JS80. The Shorkie LM PWMs were generated from PPMs (Figure 2A; Methods), whereas the Shorkie and Shorkie_Random_Init ISM maps were produced via an ISM analysis that systematically substituted each nucleotide with the three alternatives. (D) Canonical S. cerevisiae splicing motifs81. (E-G) Shorkie ISM maps for splicing motifs in DTD1, MIMS2, and two-intron gene HOP2. (H) TF-MoDISco-identified motifs on Shorkie ISM maps: curated yeast database motifs (top) and Shorkie-derived motifs (bottom).

Time-course analysis of stress-responsive transcription factor induction.
(A-E) MSN2 induction at the ATG42 promoter region (–450 to +50 bp relative to the TSS; chrII:515,214–515,714), sampled at seven timepoints blabeled in minutes. (A) Shorkie ISM sequence logos: rows correspond to successive timepoints (top to bottom), with the bottom row showing the reference. Key TF-binding motifs are annotated. (B) Experimental fold-change in reads per million (RPM) (blue) versus Shorkie-predicted signal (orange) across the ATG42 locus at each timepoint. (C) Heatmap of pairwise Euclidean distances between ISM logos, illustrating temporal divergence in motif strength and composition. (D) TF-MoDISco-identified motifs extracted from ΔT ISM matrices relative to T0. (E) Boxplot of normalized Pearson’s R between experimental and predicted profiles across all S. cerevisiae genes for MSN2 induction at each timepoint. (F–J) MSN4 induction at the TSL1 promoter region (–450 to +50 bp relative to the TSS; chrXIII:70,173–70,673), with panels analogous to (A–E).

Evaluation of Shorkie’s predictions of promoter variant effects using MPRA data.
(A) Experimental schematic showing MPRA sequences inserted at positions 100–200 bp upstream of the TSS in 10 bp increments for selected yeast genes. (B-C) Classification performance distinguishing highversus low-expression sequences across upstream insertion sites assessed by (B) AUROC and (C) AUPRC. Genes were stratified into three RNA-seq expression quantiles (5–25%, 25–75%, and 75–95%); dashed colored lines represent individual genes, and black lines depict mean ± standard error. (D–E) Comparison between Shorkie predictions (log fold-change scores) and DREAM-RNN model predictions with experimentally measured expression for (D) native yeast sequences and (E) challenging sequences. (F–H) Model performance evaluated for specific regulatory variant sets: (F) single-nucleotide variants (SNV), (G) motif perturbations, and (H) motif tiling constructs. (I) RNA-seq coverage predictions comparing Shorkie and DREAM-RNN against experimentally measured coverage.

Shorkie accurately predicts cis-eQTL variant effects.
(A) Positive eQTL example demonstrating reduced expression associated with the alternate allele at OMA1 (chrXI:603,195–604,232). (B) Positive eQTL showing increased expression associated with the alternate allele at LAP3 (chrXIV:200,569–201,933). (C) Computation of log fold-change expression scores for evaluating variant effects. (D) Generation of negative eQTL controls matched by genomic characteristics from ~1,000 natural yeast isolates. (E,F) Precision-recall (PR) curves comparing Shorkie and DREAM models for Caudal et al. (E) and Kita et al. (F) datasets. (G,H) AUPRC scores by TSS distance bins in Caudal et al. (G) and Kita et al. (H) datasets. (I–N) ISM maps centered on eQTL SNPs, highlighting regulatory motifs identified by Shorkie (SNP in light-blue) and DREAM-RNN, with adaptor sequences in gray.