Peer review process
Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.
Read more about eLife’s peer review process.Editors
- Reviewing EditorMurim ChoiSeoul National University, Seoul, Republic of Korea
- Senior EditorMurim ChoiSeoul National University, Seoul, Republic of Korea
Reviewer #1 (Public review):
Summary:
In this manuscript, Tuñí-Domínguez et al. present a large-scale, cross-study meta-analysis of 2,356 cell-free RNA sequencing (cfRNA-Seq) samples. The authors aimed to systematically evaluate the impact of pre-analytical variables and library preparation protocols on biological interpretation. By harmonizing publicly available and internally generated datasets through a uniform bioinformatics pipeline, the study seeks to establish standard quality control (QC) metrics and provide evidence-based guidelines for protocol selection in cfRNA-Seq biomarker discovery.
Strengths:
A major strength of the study is the scale and breadth of the harmonized dataset. The use of a common computational pipeline reduces variation arising from differences in bioinformatic processing and enables more direct comparisons among published datasets. The authors examine multiple complementary dimensions of data quality rather than relying on a single sequencing metric. The inclusion of variance-partition analyses, healthy-control-only analyses, and analyses restricted to samples with low gDNA contamination strengthens the evaluation of technical heterogeneity. The public availability of the analysis code and processing configurations further increases the reproducibility and potential utility of this work. The study convincingly demonstrates that technical and pre-analytical factors are major sources of variation across existing plasma cfRNA-seq datasets.
Weaknesses:
Several central conclusions are broader than the current cross-study design can fully support. Protocol category is strongly associated with dataset, laboratory, sample source, and collection procedure, making intrinsic protocol effects difficult to separate from study-specific effects, particularly for categories represented by only one or a few studies. The conclusion that technical variation generally overwhelms biological variation also requires qualification because diverse diseases and cancer types are combined into broad phenotype categories, and many phenotypes are concentrated within individual datasets. The interpretation and additional value of the proposed NG80 and NP80/NG80 metrics require further support, especially given their variable behavior across datasets. Most importantly, the thresholds used in the proposed universal QC framework are not independently validated and are strongly influenced by library-preparation strategy. The current findings therefore support protocol-dependent comparisons and identification of technical trade-offs more strongly than they support a universal definition of library quality.
Major concerns:
(1) Protocol effects remain difficult to distinguish from study-specific effects.
The manuscript interprets Broad Protocol Category as a major determinant of transcriptomic variation. However, protocol, dataset, laboratory, sample source, and collection procedure are strongly interconnected. Although both Dataset and BPC are included in the variance-partition model, several protocol categories are represented by only a limited number of studies.
This is particularly problematic for WRO, which is represented by a single study. Its apparent characteristics therefore cannot be separated from Reggiardo-specific laboratory, cohort, provider, or sample-processing effects. The authors should clarify the stability of the variance-partition results and, where feasible, provide sensitivity analyses. Alternatively, conclusions based on protocol categories represented by one or very few studies should be explicitly presented as study-specific observations rather than broadly validated protocol properties.
(2) The interpretation and robustness of the proposed library-diversity metrics require further support.
The manuscript attributes the high NG80 values in the Block and Sun datasets to gDNA contamination. However, other datasets with similarly low FSR values, including Wang and Giráldez, do not show comparably high NG80 values. This suggests that gDNA contamination alone is insufficient to explain library diversity and that sequencing depth, fragment length, mapping behavior, or library construction may also contribute.
The NP80/NG80 ratio is conceptually reasonable, but its additional value beyond the RNA-biotype composition shown in Figure 3C is unclear, particularly given the large variability observed in WRR datasets. The authors should further examine the relationships among FSR, NG80, sequencing depth, and protocol characteristics and assess the stability of NP80/NG80. If additional validation is not feasible, the interpretation and generality of these metrics should be moderated.
(3) The conclusion that technical variation generally overwhelms biological variation requires qualification.
The manuscript provides convincing evidence that technical heterogeneity is a major source of variation in the combined cross-study dataset. However, the conclusion that donor phenotype contributes only negligible variation may be broader than the analysis supports.
Phenotypes are reduced to healthy, cancer, and non-cancer disease categories despite substantial biological heterogeneity, and many disease groups are concentrated within particular studies. Consequently, disease-specific biological variation may partly be assigned to the Dataset term. A similar issue affects the cellular-origin analysis, where collection center and phenotype are substantially associated in the Chen dataset, making their individual contributions difficult to distinguish.
Where sample sizes permit, the authors should examine more specific disease categories or perform within-dataset analyses. Otherwise, the conclusions should be narrowed to state that technical variation dominates the present heterogeneous cross-study aggregation, rather than implying that phenotype-associated cfRNA signals are generally negligible.
(4) The proposed universal QC framework is insufficiently justified and protocol-dependent.
Figure 6 defines high-quality libraries using NG80 >1,000 together with FSR >20% or FER >75%. However, the manuscript does not explain how these thresholds were selected or validate them against an independent measure of reproducibility, analytical performance, or biomarker utility.
These criteria are also strongly affected by library-preparation strategy. FER and FSR favor libraries enriched for exonic or spliced RNA, whereas the NG80 cutoff disadvantages WRR libraries containing abundant noncoding transcripts. This is difficult to reconcile with the recommendation of WRR for exploratory transcriptomic and microbial analyses.
The authors should justify the threshold selection and assess its sensitivity and protocol dependence. Ideally, the criteria should be validated against an independent performance endpoint. Otherwise, Figure 6 should be reframed as a descriptive comparison, and protocol- or application-specific guidance should replace a universal binary definition of library quality.
Reviewer #2 (Public review):
Summary:
This manuscript systematically evaluates the impact of experimental workflows on plasma cell-free transcriptome sequencing (cfRNA-seq) data. The authors integrate a large number of cfRNA sequencing datasets from multiple publicly available studies and establish a unified bioinformatics framework to systematically assess the effects of experimental workflows, genomic DNA contamination, library diversity, and preanalytical factors on cfRNA transcriptomic profiles.
Strengths:
This study addresses the technical heterogeneity that may hinder cfRNA biomarker discovery and clinical translation and is of substantial value.
Weaknesses:
The effects of disease phenotype, technical confounding, criteria for library quality, and conclusions regarding DNase treatment require further clarification and validation.
Major Points:
(1) The conclusion that donor phenotype explains only a small fraction of transcriptomic variation requires further support from within-study analyses.
The authors conclude from variance partitioning across all studies that phenotype explains only a small fraction of cfRNA transcriptomic variation. However, the included studies encompass different diseases, while phenotype is simplified into healthy, cancer, and non-cancer disease categories, and the overall transcriptomic variation is strongly influenced by study-specific and experimental workflow batch effects. Therefore, the cross-study pooled analysis may underestimate genuine disease-associated cfRNA differences within individual studies conducted under the same experimental workflow.
Recommendation: The authors are encouraged to perform within-study phenotype analyses in datasets that include both healthy controls and disease samples and have sufficient sample size, and to quantitatively estimate the proportion of transcriptomic variance explained by phenotype. For example, the Zhu dataset includes healthy controls and liver cancer samples, and the authors have already observed relatively clear phenotype-associated clustering between the two groups. The contribution of healthy-versus-liver-cancer phenotype to transcriptomic variation could therefore be quantified within this dataset. If similar results are obtained across multiple independent cohorts, the findings could then be summarized across studies. The authors should also note that a low contribution of phenotype to global transcriptomic variance does not necessarily imply that disease-associated cfRNA signals lack biological or clinical relevance.
(2) Comparison of the relative contributions of technical factors and disease phenotype may be affected by confounding.
Figure 1C shows that phenotype is strongly or even completely confounded with technical variables such as collection center and centrifugation protocol in some cohorts. Nevertheless, the variance partitioning analysis across all samples is used to conclude that technical factors are the primary sources of variation, whereas phenotype contributes little. In the presence of such confounding, technical effects and disease-associated biological effects may not be reliably estimated independently, and this conclusion therefore requires more direct validation.
Recommendation: The authors are encouraged to perform an independent within-study variance analysis in cohorts in which technical variables and phenotype are relatively balanced. For example, in the Moufarrej cohort, phenotype is essentially unconfounded with collection center/centrifugation protocol (Cramer's V = 0). Phenotype and relevant technical variables could be included simultaneously in a within-cohort model to quantify their respective contributions to cfRNA transcriptomic variation. If technical factors still explain a larger fraction of variance in such relatively unconfounded cohorts, this would provide stronger support for the central conclusion of the manuscript.
(3) The use of NG80 as a criterion for defining "high-quality libraries" requires further validation.
The authors use NG80 as a metric of library diversity and further apply NG80 > 1,000 as one criterion for defining high-quality libraries in Figure 6. However, because NG80 is based on gene counts, it may be affected by sequencing depth. In addition, gDNA contamination can artificially increase NG80, whereas genuinely abundant non-coding RNAs in WRR libraries can lower NG80. Therefore, a higher NG80 does not necessarily indicate better overall library quality, and the metric may reflect both technical quality and genuine RNA composition.
Recommendation: The authors are encouraged to re-evaluate NG80 after downsampling samples to a common number of mapped fragments and to examine the relationship between NG80 and sequencing depth. The rationale for the NG80 > 1,000 threshold should also be further justified, and the impact of alternative NG80 thresholds on high-quality library classification and the main conclusions should be assessed. Unless there is evidence that this threshold reliably predicts library reproducibility or biomarker-related information content, library diversity and overall library quality should be clearly distinguished, and NG80 should not be presented as a universal criterion for high-quality libraries.
(4) Conclusions regarding DNase treatment should be interpreted more cautiously.
The Toden study did not explicitly report DNase treatment. The manuscript infers that DNase digestion was performed based on the fact that this study originated from the same laboratory as other studies and used a similar workflow; this inference should not be treated as an established experimental fact. In addition, the authors state that double DNase treatment is the most effective approach among non-EB workflows, but this conclusion is mainly based on cross-study comparisons, in which DNase strategy varies together with laboratory, sample handling, and cohort-specific factors. The current evidence is therefore insufficient to establish that double DNase treatment itself is superior.
Recommendation: The authors are encouraged to label the DNase status of the Toden study as "not reported" or "inferred", unless confirmation can be obtained from the original authors. The conclusion that double DNase treatment is the most effective approach should also be tempered, with explicit acknowledgment that it requires direct parallel validation using the same samples under different DNase treatment strategies.