Peer review process
Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.
Read more about eLife’s peer review process.Editors
- Reviewing EditorWenfeng QianChinese Academy of Sciences, Beijing, China
- Senior EditorAlan MosesUniversity of Toronto, Toronto, Canada
Reviewer #1 (Public review):
Summary:
The study identifies and characterizes a set of amino acid states that make the protein robust to other mutations, to the point of being able to compensate mutations that render wildtype proteins entirely non-functional. The study uses a previously published dataset and uses it to find and study such super-compensators. It then analyzes the biophysics and fitness landscape structure of what may be behind the compensation, identifying stability as an important parameter that, nevertheless, is not sufficient to explain all of the compensatory effect. These findings have important implications for our understanding of protein evolution, with these super-compensators possibly acting in a role of "permissive mutations" and opening up evolutionary trajectories that may be closed without them. Perhaps the identification of such super-compensator substitutions can be incorporated into various protein design approaches.
Strengths:
The paper presents a compelling case with a rigorous analysis of the expected error rates of observation. While not unique, the current state-of-the-art in the field typically does include experimental error rate estimation like this work. The paper also does a good job in exploring the issue, including looking at plausible biophysical basis of super-compensators.
Weaknesses:
The paper lacks rigor in talking about evolutionary-related issues of the state of the fitness landscape. As an example, the paper mentions that these super-compensators flatten the landscape. While I understand where this is coming from, I think that the fitness landscape in this context is a static entity and cannot be flattened or otherwise altered. A much more accurate description is that a sequence with a super-compensator is located in a flatter-than-expected segment of the fitness landscape, or on a flat fitness ridge. These issues are more semantic in nature, and while the manuscript would benefit from it being shown to an expert in molecular evolution or fitness landscapes, this issue does not take away from the importance of the results.
Reviewer #2 (Public review):
Summary:
This manuscript presents an interesting and conceptually valuable analysis of compensatory evolution using a large combinatorial deep-mutational-scanning dataset for yeast His3p.
Strengths:
I particularly like the identification of "super compensatory" substitutions that improve fitness across diverse genetic backgrounds and apparently reduce the sensitivity of the local fitness landscape to subsequent mutations. The work connects epistasis, protein stability, mutational robustness, and evolvability in a clear and potentially broadly relevant manner.
The authors provide several complementary lines of evidence in support of this central conclusion. In particular, the new experimental validation of S189A is an important strength because it directly demonstrates that a predicted super compensator can buffer the effects of diverse deleterious substitutions, while analyses of additional DMS datasets from other proteins and assay systems suggest that the phenomenon is not restricted to the original His3p landscape.
Weaknesses:
The structural analysis currently relies primarily on correlations with RSA, weighted contact number, conservation, and Rosetta-predicted changes in folding or binding energy. For super compensators, the mechanistic evidence is largely limited to predicted stabilization and individual examples, such as the proposed salt bridge between 110D and R112. I believe that the newly developed structure-aware deep-learning approaches could provide useful information on the mechanism of super compensators. For example, an inverse-folding model such as ESM-IF1 could score complete multi-mutant sequences conditioned on the His3p backbone and test whether adding a super compensator restores sequence-structure compatibility across backgrounds. More recent multimodal mutation-effect or stability models could similarly be used to cross-check the Rosetta results, including models that explicitly support combinatorial mutations. I would not recommend simply comparing AlphaFold confidence scores between mutants, because current structure predictors are not necessarily sensitive to subtle mutation-induced energetic or conformational changes.
The manuscript states that the pipeline was applied to 217 ProteinGym datasets and concludes that super compensators are broadly distributed across proteins and assays. However, this central generalization is described in only a few sentences and is largely relegated to Figure S7. The Methods do not explain which datasets contained sufficient combinatorial mutants to calculate compensatory ability or buffering, how many genotype pairs or quadruplets were available per substitution, or how differences in assay scale and library design were handled. This point requires clarification because supercompensation is inherently a background-dependent property and cannot be established from single-mutant measurements alone. ProteinGym is widely used as a substitution-effect benchmark, and many of its constituent assays primarily contain single substitutions; for example, an analysis of an earlier ProteinGym collection reported that 76 of 87 assays contained only single substitutions. It is therefore unclear how the same compensatory-interaction pipeline could be applied uniformly to all 217 datasets.
The analysis of 335 His3p orthologs in Discussion is potentially very interesting, but co-occurrence between super compensators and putatively deleterious amino-acid states does not by itself demonstrate evolutionary compensation. Closely related species share substitutions through common ancestry, and both states could be associated with a particular lineage or ecological context. A tree-aware analysis would considerably strengthen this result. The authors could reconstruct ancestral states and ask whether acquisition of a super compensator tends to precede or accompany otherwise deleterious substitutions. Alternatively, they could use phylogenetically informed permutations that preserve substitution frequencies and shared ancestry.
Reviewer #3 (Public review):
The manuscript by Jiang and co-authors presents an analysis of experimental measurements (about 400k variants) from a deep mutational scan of the HIS3 enzyme. The authors assess the ability of a genotype to be "rescued" and show that this depends on mutation sites (in particular their solvent accessibility) and mutation effects (should be mild on folding stability or binding affinity). They further identify a set of super-compensatory mutations, and their results suggest that these mutations flatten the fitness landscape.
This finding is interesting and likely of interest to a broad community. The analysis seems sound.
However, I have a number of major concerns regarding the presentation and positioning of the work.
(1) It would improve the manuscript to clarify the present contribution with respect to a previous study by the same authors, namely Pokusaeva et al. 2019. Did the authors apply the same protocol to generate a new library of mutants, or did they re-analyse an already published library? If the library is not new, ambiguous sentences like "Nevertheless, to our knowledge, the His3p library remains one of the largest and most comprehensive resources that contains multi-site mutants" should be reformulated.
(2) Pokusaeva et al. 2019 is cited for the library and also for the deep neural network. It would be beneficial to briefly describe the architecture, the inputs and outputs, and the training procedure. Was the network trained on the current library? What is the purpose of this network? It looks more like an additive linear model (except for the global sigmoid) than a deep neural network. How does it relate to global epistasis models? The sigmoid function is designed to capture plateauing effects; doesn't that introduce some circularity issue in the reasoning?
(3) Are the super-compensatory mutations observed (conserved) across evolution? Beyond the fact that they are accompanied by mildly deleterious mutations in natural sequences. Can we predict them with variant effect predictors?
(4) The AAindex mention should be accompanied by a citation.
(5) Equations should be numbered. WCN formula seems to contain misformatting issues.
(6) A more explicit description of the structural data analysed (which PDB entry?) should be provided.
(7) I believe the citation Van Cleve and Weissman 2015 for the ProteinGym benchmark is incorrect. Additionally, is the Rosetta citation adequate?
(8) How is the definition of rescueability sensitive to the threshold choice?