Overview of datasets and processing pipeline.

Sample batches identified within some studies are depicted separately but grouped together using horizontal brackets below the X axis. “Unspecified”: not reported in the publication. (A) Summary of pre-analytical variables across datasets. Datasets were split into distinct batches when those differed by at least one pre-analytical variable. Note that the Toden dataset is annotated per their published Methods, which do not mention DNase treatment. As discussed later in the main text, we speculate that this likely reflects an unintentional omission from the authors, rather than a genuine absence of treatment. (B) Overview of datasets under consideration, broken down by donor phenotype. (C) Association between various pre-analytical variables and donor phenotype across datasets. Cramér’s V scores are reported for each variable. Variables strongly confounding phenotype approach a score of 1. “Uniform” means that the variable of interest is constant within the dataset. Only datasets studying more than one phenotypic condition were included. For clarity, we also excluded pre-analytical variables that were fully uniform or unspecified across all datasets. Those were: “Blood collection tube”, “RNA extraction kit”, “DNase treatment”, “Library prep kit”, “Library selection” and “cDNA library type”. *One sample with missing information from the Toden dataset was removed to allow the calculation. **17 samples with missing phenotype information from the Ibarra (plasma) dataset were removed. (D) Illustration of the nf-core/rnaseq Nextflow pipeline used for the uniform processing of RNA-seq data.

Microbial reads in cfRNA-Seq libraries.

(A) Fraction of sequencing reads mapping to the human genome across datasets. Each library is represented by a dot. The distribution of the metric within a given dataset is represented by a boxplot. (B) Taxonomic profile of dataset batches. Relative abundance values were averaged across samples, grouped into the simplified categories: human (Homo sapiens), fungi (Fungi), other_eukaryotes (Eukaryota except human and Fungi), bacteria (Bacteria), other (Archaea, Viruses, and other unspecific taxa), and unclassified.

gDNA contamination and library diversity in cfRNA-Seq libraries.

(A) Fraction of spliced reads. (B) Library diversity (NG80). (C) Mean RNA biotype representation. (D) NP80/NG80 ratio in gDNA-free datasets.

Cellular origin of the cell-free transcriptome.

Only gDNA-free datasets are represented (A) Average predicted cell type contribution to cfRNA. (B) Clustered heatmap showing a cell type fraction batch effect in the Chen dataset. Only major cell types (MCTs) are represented on the Y axis. Heatmap cells are colored according to the corresponding predicted cell type contribution to the cfRNA library (X axis), in relative abundance. Each library is annotated with its corresponding relevant metadata (“phenotype” and “collection_center”, top rows)

Determinants of cell-free RNA gene abundance profiles.

(A-F) Scatterplot visualizations of gene abundance-based principal components 1 and 2. Each dot corresponds to a sequencing library. (A-C): All datasets. (E-F) gDNA-free datasets only. In this latter representation, note that the left and right clusters roughly correspond to EB and WRR/WRO datasets, respectively (Supplementary fig. 7I). (A) Colored by dataset/batch label. (B) Colored by BPC (Broad Protocol Category). (C) Colored by log(NG80). (D) Colored by fraction of spliced reads (FSR). (E) Colored by fraction of platelet RNA. (F) Colored by fraction of microbial reads. (G) Variance Partition Analysis (VPA). For each variable of interest (X axis), a genome-wide violin plot of the distribution of variance explained by this variable across all genes is shown. Variables are sorted in descending order by median variance explained across genes, except the residuals, which are represented on the right by convention.

Fraction of high-quality libraries in each dataset.

High-quality libraries are defined as those exhibiting an NG80 of more than 1,000 and an FSR greater than 20% or an FER of more than 75%. A more detailed sample survival analysis is available in Supplementary fig. 9.