Protein sequence generation with evolutionary diffusion.

(A) (Left) Evolution has sampled a tiny fraction of possible protein sequences. Experimental structures have been determined for even fewer proteins. (Right) EvoDiff is a generative discrete diffusion model trained on natural protein sequences. Sampling from EvoDiff yields new protein sequences that may perform desired functions. (B) Discrete diffusion models consist of controlled corruption and learned denoising processes. In the masked corruption process, input tokens are masked in an order-agnostic fashion (bottom, left). In the discrete corruption process, inputs are corrupted via a Markov process controlled by a transition matrix capturing amino acid mutation frequencies (bottom, right). (C) EvoDiff enables unconditional generation of protein sequences or MSAs. Starting from masked or uniformly sampled inputs xT, EvoDiff generates new sequences or MSAs by reversing the corruption process, iteratively denoising xt into realistic sequences or MSAs x0. (D) Controllable protein design with EvoDiff, via conditioning on evolutionary information encoded in MSAs (left); inpainting functional domains from masked portions of a sequence (middle); or scaffolding structural motifs without explicit structural information (right). thermore, EvoDiff-OADM is the only model variant where performance scales with increased model size (Table S1; Fig. S1).

EvoDiff generates realistic and structurally-plausible protein sequences.

(A) Workflow for evaluating the foldability and self-consistency of sequences generated by EvoDiff sequence models. (B-C) Distributions of foldability, measured by sequence pLDDT of predicted structures (B), and self-consistency, measured by scPerplexity (C), for sequences from the test set, EvoDiff models, and baselines (n=1000 sequences per model; box plots show median and interquartile range). (D) Sequence pLDDT versus scPerplexity for sequences from the test set (grey, n=1000) and the 640M-parameter OADM model EvoDiff-Seq (blue, n=1000). (E) Predicted structures and metrics for successfully expressed and characterized unconditional generations from EvoDiff-Seq, the 640M-parameter OADM model. OmegaFold predictions, colored by pLDDT, are shown, and the average pLDDT for each structure is reported. % coverage and % identity to the top BLAST hit is denoted below each design. (F) Circular dichroism (CD) spectra for designed sequences from (E). (G) Structural composition of each sequence design as inferred from CD spectra (blue) versus from OmegaFold (grey). AlphaFold predictions are included in Fig. S6 for comparison.

Generated protein sequences capture natural distributions of protein functional and structural features.

(A) UMAP of ProtT5 embeddings, annotated with FPD, of natural sequences from the test set (grey, n =l000) and of generated sequences from EvoDiff-Seq (blue, n =l000) and ESM-2 (red, n =l000), and inferred sequences inverse-folded from structures from RFdiffusion (orange, n=l000). (B) Multivariate distributions of helix and strand structural features in generated sequences, based on DSSP 3-state predictions (n =l000 samples from each model or the validation set) and annotated with the Kullback-Leibler (KL) divergence relative to the test set.

EvoDiff-MSA enables evolution-guided sequence generation.

(A) A new sequence is generated from EvoDiff-MSA via diffusion over only the query component. Generations are evaluated for diversity and self-consistency and for the quality and consistency of their predicted structures. (B-E) Distributions of pLDDT (B), scPerplexity (C), sequence similarity (D; dashed line at 25%), and TM-score (E; dashed line at 0.5) for sequences from the validation set, EvoDiff-MSA, ESM-MSA, and a Potts model (n=250 sequences per model; box plots show median and interquartile range). (F) Sequence pLDDT versus scPerplexity for sequences from the validation set (grey, n=250) and EvoDiff-MSA (blue, n=250). (G) Predicted structures and metrics for structurally plausible generations from EvoDiff-MSA.

EvoDiff generates intrinsically disordered regions.

(A) A new IDR sequence is generated from EvoDiff-Seq or EvoDiff-MSA by inpainting disordered residues in the query sequence. DR-BERT is then used to predict disorder scores for the original and regenerated sequences. (B) Distributions of disorder scores over disordered and structured regions for sequences with true (grey), inpainted (blue), and randomly-sampled (red) IDRs (n=100 sequences per condition; box plots show median and interquartile range). (C) Distribution of sequence similarity relative to the original IDR for generated IDRs from EvoDiff-Seq (blue, dashed) and EvoDiff-MSA (blue, solid) (n=100; dashed line at 25%). (D) Structure of the yeast Cox15 protein, with the disordered N-terminal mitochondrial targeting signal highlighted in yellow. EvoDiff is used to inpaint this disordered region. (E) Microscopy assay in yeast to assess cellular localization of Cox15 proteins with EvoDiff-designed targeting signals. Mitochondrial localization is measured by quantifying the fluorescence overlap between endogenous Cox15 (mScarlet) and designed Cox15 (sfGFP). (F) Fluorescence microscopy imaging of yeast cells. Columns from left to right show: the N-terminal IDR sequence used for a given Cox15-sfGFP design (rows); signal from the designed Cox15 (sfGFP, colored in green); signal from the endogenous Cox15 (mScarlet, colored in red); overlaid images. All images show brightfield in grey, to visualize cell boundaries. Images for all 8 designs are included in Fig. S15 and S16. Scale bars = 8 μm. (G) Per-cell pixel-wise Pearson correlation coefficients of signal intensities between the mScarlet channel and the sfGFP channel over n=183-313 cells segmented by YeastSpotter for each design. Correlations for a biological replicate are shown in Fig. S17.

EvoDiff scaffolds functional motifs without explicit structural information.

(A) Number and identity of successful scaffolds from n=100 trials for EvoDiff-Seq (x-axis) versus EvoDiff-MSA (y-axis) across scaffolding problems in which at least one method succeeds. (B) Performance comparison of sequence-based scaffolding via EvoDiff-MSA (y-axis) versus structure-based scaffolding via RFdiffusion (x-axis) across scaffolding problems in which at least one method succeeds (n=100 trials per problem). (C) Distributions of TM-scores of successfully generated scaffolds from EvoDiff models relative to the true structures (dashed line at 0.5; box plots show median and interquartile range). (D) EvoDiff-Seq successfully scaffolds the binding site of compact calmodulin. The generated sequence, OmegaFold-predicted structure (colored by pLDDT), and computed metrics for EvoDiff-Seq-1PRW-Design1 are shown. Scaffolding motif is shown in green, and the structure and sequence of calmodulin (1PRW) are shown in dark grey. (E) Background-corrected absorbance spectra of the native 1PRW (grey), EvoDiff-Seq-1PRW-Design-1 (blue), and control buffers (red) with (solid line) and without (dashed lines) Ca2+. (F) CD spectra of native 1PRW (grey) and EvoDiff-Seq-1PRW-Design-1 (blue). (G) EvoDiff successfully scaffolds the transactivation domain of p53. Generated sequence, OmegaFold-predicted structure (colored by pLDDT), and computed metrics for EvoDiff-MSA-p53-Design-3, are shown. The transactivation domain is shown in green, and the AlphaFold-predicted structure and sequence of p53 are shown in dark grey. (H) BLI measurements for EvoDiff-MSA-p53-Design-3 measured over 5 concentrations. (I) In vitro binding affinities of all EvoDiff designs against the full-length human MDM2. Measurements for all designs were performed in triplicate (Fig. S20-S24).

Perplexity as a function of corruption step for EvoDiff sequence models.

Test-set perplexities at sampled intervals of the degree of corruption, specifically the diffusion timestep for D3PM models, the fraction of masked residues for OADM and masked language models, and the fraction of evaluated sequence for LRAR models. Intervals reflect evenly spaced windows of 50 timesteps for D3PM models or 10% masking for masked models.

Perplexity as a function of corruption step for EvoDiff MSA models.

Test-set MSA perplexities at sampled intervals of the degree of corruption, specifically the diffusion timestep for D3PM models and the fraction of masked residues for OADM and ESM models. The test-set evaluated for each model was sampled using the same sampling scheme assigned during training. Intervals reflect evenly spaced windows of 50 timesteps for D3PM models or 10% masking for masked models.

Summary statistics for structural plausibility metrics for sequence models.

(A-B) Distribution of pLDDT and scPerplexity metrics for sequences from the test set, 38M parameter EvoDiff and baseline models (A), and 640M parameter EvoDiff and baseline models (B) (n=l000 sequences per model). Test and Random baselines are reproduced in (A) and (B) for reference.

Sequence pLDDT versus self-consistency perplexity for EvoDiff sequence models.

(A-B) Results for sequences from 38M parameter EvoDiff, baseline models, and test data plotted alone (A), and 640M parameter EvoDiff and baseline models (B) (various colors, n=1000), except for EvoDiff-OADM-640M (EvoDiff-Seq, shown in Fig. 2C), relative to sequences from the test set (grey, n=1000). The self-consistency perplexity (ESM-IF Perplexity) is computed using sequences inverse-folded by ESM-IF.

Representative proteins generated by EvoDiff-Seq.

(A) Predicted structures and metrics for representative structurally plausible generations from EvoDiff-Seq, the 640M-parameter OADM model. (B) Corresponding per-residue pLDDT for pLDDT scores computed based on the OmegaFold predicted structures, for individual residues in representative high-fidelity generations from EvoDiff-Seq (Fig. 2D). Points are colored by pLDDT (0-100, red to blue).

Comparison of structure-prediction methods for unconditional EvoDiff-Seq designs

(A) AlphaFold predictions (colored by pLDDT) overlaid on OmegaFold (grey). The AlphaFold pLDDT, and backbone RMSD (bbRMSD) taken between the two structures is denoted. (B) Percent composition extracted from CD spectra (blue) compared to structure prediction methods (AlphaFold and OmegaFold, grey).

Coverage of sequence and functional space for generated distributions from 38M parameter EvoDiff sequence models and baselines.

UMAP of ProtT5 embeddings, annotated with FPD, of natural sequences from test set (grey, n= l000 plotted) and of generated sequences from EvoDiff 38M parameter models and baselines (various colors, n=l000).

Coverage of sequence and functional space for generated distributions from 640M parameter EvoDiff sequence models and baselines.

UMAP of ProtT5 embeddings, annotated with FPD, of natural sequences from test set (grey, n= 1000) and of generated sequences from EvoDiff 640M parameter models and baselines (various colors, n=l000). A visualization of sequences from the validation set (dark grey, n= 1000) is included for reference. The visualization for the 640M OADM model is excluded due to inclusion in Fig. 3A.

Structural features in generated sequences from all sequence models.

(A-B) Multivariate distributions of helix and strand features in sequences from 38M (A) and 640M (B) parameter models, and baselines based on DSSP 3-state predictions and annotated with the KL divergence relative to the test set (n=1000 samples from each model). In (B), the distribution for the 640M OADM model is excluded (see Fig. 3B); the distribution for random sequences (n= 1000) is provided as reference.

Sequence pLDDT versus scPerplexity for EvoDiff MSA models,

for sequences from the validation set (grey, n=250) and evaluated MSA models (various colors, n=250), except for EvoDiff-OADM-MSA-Max (EvoDiff-MSA, shown in Fig. 4F).

Baseline performance of DR-BERT evaluator.

(A-C) Distributions of DR-BERT predicted disorder scores across disordered and structured regions for sequences with true (A), scrambled (B), and randomly sampled (C) IDRs (n=l00).

Performance of DR-BERT evaluator on disorder regions.

(A-B) Disorder scores predicted by DR-BERT for true (x-axis) vs. generated (y-axis) IDRs for the same given sequence, for generations from EvoDiff-Seq (A) and EvoDiff-MSA (B). Each dot represents an individual IDR (n=l00). The Pearson R is given for each of EvoDiff-Seq and EvoDiff-MSA.

Examples of predicted disorder scores

and corresponding sequences for repre-sentative generated (top row) and native (bottom row) IDRs from EvoDiff-Seq (A) and EvoDiff-MSA (B).

In-silico predictions for inpainted Cox15 designs

(A) Disorder scores predicted by DR-BERT and (B) MitoFates probability of presequences for 600 EvoDiff designs (blue) and 600 randomly sampled sequences (red). Scores for the native Cox15 IDR are denoted with a dashed yellow line. (C) Disorder and (D) Mitofates probability of presequence predictions for tested Cox15 IDRs, including EvoDiff designs (blue), native Cox15 IDR (yellow), and N1 (red).

Fluorescence microscopy imaging of yeast cells for EvoDiff-Seq generated Designs.

Columns from left to right show: the N-terminal IDR sequence used for a given Cox15-sfGFP design (rows); signal from the designed Cox15 (sfGFP. colored in green); signal from the endogenous Cox 15 (mScarlet, colored in red); overlaid images. All images show brightfield in grey, to visualize cell boundaries. Images for the native IDR, N1, and Δ2-45 are included for reference.

Fluorescence microscopy imaging of yeast cells for EvoDiff-MSA generated Designs.

Columns from left to right show: the N-terminal IDR sequence used for a given Cox15-sfGFP design (rows); signal from the designed Cox15 (sfGFP, colored in green); signal from the endogenous Cox15 (mScarlet, colored in red); overlaid images. All images show brightfield in grey, to visualize cell boundaries. Images for the native IDR, N1, and Δ2-45 are included for reference.

Per-cell pixel-wise Pearson correlation coefficients of fluorescence microscopy signal intensities

between the mScarlet channel and the sfGFP channel. Independent replicates are measured over n=110-264 cells segmented by YeastSpotter for each design. The native IDR is shown in yellow, N1 in red, and EvoDiff designs in blue.

EvoDiff-Seq and EvoDiff-MSA scaffold binding domains in sequence space

(A-D) Generated sequences, predicted structures, and computed metrics for representative scaf-folding examples from EvoDiff-Seq (A/B) and EvoDiff-MSA (C/D). Motif is shown in green, original scaffold in gray, and generated scaffold in blue.

Calmodulin binding assay.

Included is the native IPRW protein (grey), successful EvoDiff design (EvoDiff-Seq-lPRW-Design-l, blue), an unsuccessful EvoDiff design (EvoDiff-Seq-lPRW-Design-2, orange), and buffer (red), with (solid-line) and without (dashed-line) Ca2+. The mean value is reported, and error bars are represented by the std. dev taken from three independent trials

BLI results for positive controls.

BLI results for successful MDM2 binders generated by EvoDiff-Seq.

BLI results for successful MDM2 binders generated by EvoDiff-MSA.

BLI results for unsuccessful MDM2 binders EvoDiff-Seq-p53-Design-5 through EvoDiff-Seq-p53-Design-12

BLI results for unsuccessful MDM2 binders EvoDiff-Seq-p53-Design-13 through EvoDiff-MSA-p53-Design-5

Details of EvoDiff-D3PM corruption schemes.

The top and bottom rows correspond to EvoDiff-D3PM-Uniform and EvoDiff-D3PM-BLOSUM, respectively. (A) Visualization of EvoDiff-D3PM transition matrices. (B) Evolution of the number of mutations accrued as a function of the diffusion timestep t for a sample input. (C) Evolution of β as a function of the diffusion timestep t. (D) Evolution of DKL [q(xt|x0)‖p(xT)] as a function of the diffusion timestep t, indicating convergence to a uniform stationary distribution at t = 500 as DKL approaches zero.|

Performance of EvoDiff sequence models.

The reconstruction KL (Recon KL) was calculated between the distribution of amino acids in the test set and in generated samples (n=1000). The perplexity was computed on 25k samples from the test set. The minimum Hamming distance to any train sequence of the same length (Hamming) is reported for each model as the mean ± standard deviation over the generated samples. Fréchet ProtT5 distance (FPD) was calculated between the test set and generated samples. The secondary structure KL (SS KL) was calculated between the means of the predicted secondary structures of the test and generated samples.

Validation-set perplexities for EvoDiff MSA models.

The perplexity is calculated based on the ability of each model to reconstruct a subsampled MSA from the validation set. “Max Perplexity” and “Rand. Perplexity” indicate MaxHamming and Random subsampling, respectively, for construction of the validation MSA.

Structural plausibility metrics for EvoDiff sequence models and baselines.

Metrics are reported as the mean ± standard deviation for 1000 generated samples for each model.

Performance of EvoDiff MSA models in generating query sequences conditioned on MSAs.

Metrics are reported as the mean ± standard deviation over 250 generated samples for each model.

Scaffolding performance of EvoDiff-Seq.

Number of scaffolding successes out of 100 generations for RFdiffusion, EvoDiff-Seq, the LRAR baseline, the CARP baseline, and randomly sampled scaffolds (Random), for each of 17 scaffolding problems. The bottom row contains the total number of successful scaffolds generated per model.

Scaffolding performance of EvoDiff-MSA.

Number of scaffolding successes out of 100 generations for RFdiffusion, EvoDiff-MSA (Max), EvoDiff-MSA (Random), and the ESM-MSA baseline, for each of 17 scaffolding problems. The bottom row contains the total number of successful scaffolds generated per model.