Peer review process
Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.
Read more about eLife’s peer review process.Editors
- Reviewing EditorPaschalis KratsiosUniversity of Chicago, Chicago, United States of America
- Senior EditorKathryn CheahUniversity of Hong Kong, Hong Kong, Hong Kong
Reviewer #1 (Public review):
Summary:
Duan, Li, Kulkarni et al. apply a multiplexed single-cell overexpression screen (Reprogram-Seq) to combinatorially perturb 105 transcription factors across 7 target cell types in mouse embryonic fibroblasts, generating a resource of ~200,000 single-cell transcriptomes spanning over 1,300 TF combinations. They develop a framework for classifying pairwise TF-TF interactions, identify a modular, shared architecture of gene regulatory programs across diverse TF combinations, and use these tools to nominate and partially validate new reprogramming cocktails.
Strengths:
The scale of the combinatorial screen is substantial, and the resulting dataset is a genuine resource for the field. The TF-TF interaction typing framework is a useful conceptual extension of prior genetic-interaction approaches to an overexpression/reprogramming context, and the modularity finding that diverse TF combinations converge on shared gene programs is a compelling organizing principle. The authors are, for the most part, careful and appropriately hedged in their claims; the overclaiming we flag below is the exception, not the rule. We also note that the core Reprogram-Seq assay itself builds directly on the authors' own prior work; the novelty here rests on scale, the interaction framework, and the modularity analysis.
Weaknesses:
Most of the concerns raised below relate to how existing data are quantified, cited, and reconciled with the text, rather than to the underlying experiments themselves. Several quantitative and comparative claims in the Results are not fully supported by the figures cited, and some conclusions are in tension with the authors' own data. Key methodological details relevant to interpreting the central TF-TF interaction framework, including TF expression dosage and within-combination transcriptional variability, are not reported or controlled for, which limits confidence in the resulting interaction classifications. The relationship between TF number and reprogramming efficiency is not clearly distinguished from a simple combinatorial coverage effect and does not consistently generalize across batches. Experimental validation of predicted cocktails is limited to a single target cell type. The manuscript would also benefit from addressing whether TF overexpression in fibroblasts can fully capture a factor's endogenous regulatory network, given that pioneer activity and chromatin accessibility are not addressed.
Reviewer #2 (Public review):
Summary:
This manuscript presents a large-scale combinatorial transcription factor overexpression screen in mouse embryonic fibroblasts coupled with single-cell RNA-seq to systematically map the relationship between TF combinations, gene regulatory networks, and transcriptional reprogramming. Using approximately 100 transcription factors, the authors identify TF combinations that shift cells toward diverse transcriptional states, organize TF combinations into "perturbation clusters" with shared transcriptional outputs, infer modular gene regulatory networks, model pairwise TF interactions, and use pseudotime analyses to nominate candidate reprogramming TF combinations. The resulting dataset represents a potentially valuable resource for studying combinatorial TF activity and transcriptional reprogramming.
Strengths:
The primary strength of the study is its experimental scale and the breadth of the generated dataset. The Reprogram-Seq platform enables systematic interrogation of thousands of TF combinations that would be difficult to test individually, and the authors develop several computational frameworks to organize these data and generate biological hypotheses. The epicardial reprogramming analyses, including independent qPCR and protein localization experiments, provide proof-of-principle that the platform can recover biologically relevant TF combinations.
Weaknesses:
Many of the manuscript's central conclusions rely on a complex computational pipeline that is not sufficiently justified or independently validated. Identification of transcriptionally reprogrammed cells depends on co-embedding with reference atlases, yet the robustness of this analysis and the interpretation of cells occupying primary-cell clusters are not explored in depth. Similarly, the conclusions regarding modular gene regulatory networks depend critically on the perturbation clusters defined by MDE embedding and HDBSCAN clustering. These perturbation clusters form the basis for nearly all downstream analyses, including differential expression, gene module identification, gene specificity, and TF modularity, yet little evidence is provided that the clusters are robust to alternative embedding strategies, clustering parameters, or resampling approaches.
The manuscript also provides relatively limited orthogonal validation of its computational predictions. Although the epicardial analyses are validated experimentally, comparable validation is not performed for most other predicted cell fates, TF interaction classes, or perturbation modules. Consequently, many conclusions regarding the generality of modular TF activity, TF cooperativity, and the predicted reprogramming cocktails remain supported primarily by computational inference.
In addition, several aspects of the analytical workflow-including the quality filtering of TF combinations, interpretation of unclustered perturbations, selection of genes for downstream visualization, and robustness of pseudotime analyses across lineages-would benefit from greater methodological transparency.
Overall, this work provides a valuable dataset and introduces analytical approaches that will likely be useful to the community. However, in its current form, I believe the strongest biological conclusions are insufficiently validated. Additional analyses demonstrating the robustness of the computational framework, together with broader orthogonal validation of representative predictions, would substantially strengthen confidence in the proposed principles governing combinatorial transcription factor activity and transcriptional reprogramming.
Reviewer #3 (Public review):
Summary:
In this manuscript, Duan et al perform a combinatorial TF overexpression screen combined with single-cell RNA-seq (Reprogram-Seq) to extract general principles of how combinatorial TF interactions drive distinct gene regulatory networks in reprogrammed cells. Using a library of 105 TFs, they induce different cell fates, many of which resemble in vivo cell identities. They observe that combinations of TFs have better reprogramming results than inductions driven by a single TF. By looking at gene expression enrichment/depletion in different reprogrammed cell clusters, they infer functional GRNs induced by specific TF combinations and identify GRNs specific for certain cell types. They also identify TFs that could improve known TF cocktails for the induction of certain cell fates. They observe that TFs with cooperative interactions regarding the regulation of gene expression may lead to better reprogramming results, and finally, they build a bottom-up approach that utilizes the single-cell transcriptomes to predict TFs driving certain reprogrammed fates.
Strengths:
Reprogram-Seq is not new, but the strength of the study lies in the fact that the authors assess the induction outcome from a large number of different TF combinations. The authors are thus able to make broad observations, such as the modularity of GRNs and TF cooperativity, as well as propose new TFs and TF interactions to be tested for the induction of certain cell fates. The manuscript is well written, and the conclusions are, in general, supported well by the authors' analyses and data.
Weaknesses:
The study would benefit from some further analysis and discussion to better tighten the conclusions:
While both expression enrichment and depletion were used to define perturbation clusters, the authors then focused on analyzing functional gene groups only for the enriched genes. Are there any functional relations between the repressed genes within a perturbation cluster? Do the authors observe the same modularity (in terms of regulation by TFs) for repressed genes as they do for induced genes?
Can the authors give a description of how neomorphic TF interactions work? How would gene expression be affected in single vs double perturbation in those cases?
What does it mean functionally when a TF pair shows more than one type of interaction (as shown in Supplementary Table 6), and how do such interactions correlate with successful transcriptional reprogramming?
In the last Results section, the authors are able to use the transcriptome to predict the TF that was used for the induction. Could the authors discuss some plausible applications of this TF prediction method? For example, could they use it on in vivo single-cell RNA-seq of a certain cell type to predict candidate TFs for the induction of that cell type?
It would help the reader if, at the end of each Results section, the authors add a concluding paragraph, highlighting the most important conclusions and findings (same as they have done in the section titled "Combinatorial TF over-expression reprograms MEFs to diverse states").