Benchmarking biochemical networks generated by large language models
Figures
Signaling networks generated by general-purpose large language models (LLMs).
(A) Schematic of the pipeline for LLM-generated models of signaling networks. (B) Network reactions recalled by three LLMs (Gemini 3.0, orange; GPT 5.2, blue; Claude 4.6, green) compared with a ‘Ground Truth’ literature-curated and validated cardiomyocyte hypertrophy signaling network (Ryall et al., 2012; gray reactions). Edge color intensities indicate how frequently each reaction was predicted among the replicates (n=10). LLM networks were generated using iterative prompts based on the gene set of the Ground Truth hypertrophy network. (C–E) Summary of reaction recall for three literature-curated signaling networks (C, hypertrophy Ryall et al., 2012; D, fibroblast Zeigler et al., 2016; and E, mechanosignaling Tan et al., 2017) by Gemini, GPT, and Claude. * Indicates P<10–9 in one-sample t-test between LLM-generated replicates and the ground truth network.
Visualization of large language model (LLM)-generated fibroblast signaling networks, as recalled by three general-purpose LLMs.
Network reactions recalled by three LLMs (Gemini 3.0, orange; GPT 5.2, blue; Claude 4.6, green) compared with a ‘Ground Truth’ literature-curated and validated fibroblast signaling network (gray reactions). LLM-generated networks used prompts based on the gene set of the Ground Truth fibroblast network. This visualization corresponds to the analyses in Figure 1D.
Visualization of large language model (LLM)-generated mechanosignaling networks, as recalled by three general-purpose LLMs.
Network reactions recalled by three LLMs (Gemini 3.0, orange; GPT 5.2, blue; Claude 4.6, green) compared with a ‘Ground Truth’ literature-curated and validated mechanosignaling network (gray reactions). LLM-generated networks used prompts based on the gene set of the Ground Truth mechanosignaling network. This visualization corresponds to the analyses in Figure 1E.
Confusion matrices, calibration analysis, and union of predicted reactions across replicates for the large language model (LLM)-generated hypertrophy networks.
(A) Confusion matrices summarizing reaction-level prediction performance for each LLM and a null predictor (hypothetical network with no connections) relative to the manually curated hypertrophy network. Actual positives were defined as reactions present in the manually curated network, whereas actual negatives were defined as possible node-to-node reactions absent from the manually curated network. (B) Calibration analyses were performed using replicate predictions for each LLM (n=10), with confidence defined as the fraction of replicates in which a reaction was predicted. (C) Venn diagram showing the overlap between the manually curated ground-truth reactions and the union of reactions predicted by at least one replicate from each LLM (confidence ≥0.1). TP, true positive; FP, false positive; TN, true negative; FN, false negative; NPV, negative predictive value; F1, F1 score.
Confusion matrices, calibration analysis, and union of predicted reactions across replicates for the large language model (LLM)-generated fibroblast networks.
(A) Confusion matrices summarizing reaction-level prediction performance for each LLM and a null predictor relative to the manually curated fibroblast network. Actual positives were defined as reactions present in the manually curated network, whereas actual negatives were defined as possible node-to-node reactions absent from the manually curated network. (B) Calibration analyses were performed using replicate predictions for each LLM (n=10), with confidence defined as the fraction of replicates in which a reaction was predicted. (C) Venn diagram showing the overlap between the manually curated ground-truth reactions and the union of reactions predicted by at least one replicate from each LLM (confidence ≥0.1). TP, true positive; FP, false positive; TN, true negative; FN, false negative; NPV, negative predictive value; F1, F1 score.
Confusion matrices, calibration analysis, and union of predicted reactions across replicates for the large language model (LLM)-generated mechanosignaling networks.
(A) Confusion matrices summarizing reaction-level prediction performance for each LLM and a null predictor relative to the manually curated mechanosignaling network. Actual positives were defined as reactions present in the manually curated network, whereas actual negatives were defined as possible node-to-node reactions absent from the manually curated network. (B) Calibration analyses were performed using replicate predictions for each LLM (n=10), with confidence defined as the fraction of replicates in which a reaction was predicted. (C) Venn diagram showing the overlap between the manually curated ground-truth reactions and the union of reactions predicted by at least one replicate from each LLM (confidence ≥0.1). TP, true positive; FP, false positive; TN, true negative; FN, false negative; NPV, negative predictive value; F1, F1 score.
Experimental validation of perturbation responses predicted by large language model (LLM)-generated signaling network models.
(A) Representative validations of network models generated by manual curation or by LLMs (Gemini, GPT, Claude), in comparison to experiments in conditions of Angiotensin II (AngII) or isoproterenol (ISO) from the literature (Ryall et al., 2012). (B–D) Summary of systematic validations of manually curated (Ground Truth) and LLM-generated reconstructions of hypertrophy, fibroblast, and mechanosignaling network models against perturbation experiments from the literature (n=114, 83, and 171 experiments, respectively). * Indicates p<10–11 in one-sample t-test between LLM-generated model validation scores (n=10 replicates) and ground truth model validation accuracy.
Experimental validation of perturbation responses predicted by full large language model (LLM)-generated signaling network models.
(A–C) Systematic validation of manually curated ground-truth models and full, unrefined LLM-generated models for the hypertrophy, fibroblast, and mechanosignaling networks against perturbation experiments from the literature (n=114, 83, and 171 experiments, respectively). Asterisks indicate p<1 × 10⁻⁷ by one-sample t-test comparing LLM-generated model validation scores across replicates (n=10 per LLM) against the corresponding ground-truth model validation accuracy. (D) Average reaction counts in the full LLM-generated network models (n=10 per network per LLM) compared with corresponding manually curated networks.
Metabolic network models generated by general-purpose large language models (LLMs) and substrate utilization analysis.
(A) Schematic of pipeline for LLM-generated models of the core E. coli metabolic network. (B) Network reactions recalled by three LLMs (Gemini, orange; GPT, blue; Claude, green) compared with a ‘Ground Truth’ manually curated metabolic network. Bar charts illustrate how frequently each reaction was predicted by the different LLMs in the replicates (n=10). Inset shows zoomed-in model coverage of the ADK1 reaction. LLM networks were generated using iterative prompts based on the GPR gene list of the core E. coli metabolic network. (C) Summary of reaction recall accuracy by Gemini, GPT, and Claude. (D) Heatmap illustrating substrate utilization predictions of the ground truth model to and LLM-generated models compared to experimental data. Color intensity within the LLM columns indicates frequency of growth prediction among the 10 replicates. (E) Summary of model performance compared to the experimental data across carbon sources. * Indicates p<10–4 in one-sample t-test between LLM-generated replicates (n=10) and the ground truth network. Within the network visualization, diamonds indicate reactions while rectangles indicate metabolites.
Tables
Proportion of stimulatory and inhibitory connections present in each ground truth model and in the sets of connections predicted by each LLM (Claude, GPT, and Gemini).
| Hypertrophy | stimulation | inhibition |
|---|---|---|
| Ground Truth | 86.39 | 13.61 |
| Claude | 91.69 | 8.31 |
| GPT | 89.31 | 10.69 |
| Gemini | 90.54 | 9.46 |