Benchmarking biochemical networks generated by large language models

  1. Jeevan Tewari
  2. Benjamin W Dahl
  3. B Adam Bates
  4. Jason A Papin
  5. Jeffrey J Saucerman  Is a corresponding author
  1. Department of Biomedical Engineering, University of Virginia, United States
3 figures, 1 table and 1 additional file

Figures

Figure 1 with 5 supplements
Signaling networks generated by general-purpose large language models (LLMs).

(A) Schematic of the pipeline for LLM-generated models of signaling networks. (B) Network reactions recalled by three LLMs (Gemini 3.0, orange; GPT 5.2, blue; Claude 4.6, green) compared with a ‘Ground Truth’ literature-curated and validated cardiomyocyte hypertrophy signaling network (Ryall et al., 2012; gray reactions). Edge color intensities indicate how frequently each reaction was predicted among the replicates (n=10). LLM networks were generated using iterative prompts based on the gene set of the Ground Truth hypertrophy network. (C–E) Summary of reaction recall for three literature-curated signaling networks (C, hypertrophy Ryall et al., 2012; D, fibroblast Zeigler et al., 2016; and E, mechanosignaling Tan et al., 2017) by Gemini, GPT, and Claude. * Indicates P<10–9 in one-sample t-test between LLM-generated replicates and the ground truth network.

Figure 1—figure supplement 1
Visualization of large language model (LLM)-generated fibroblast signaling networks, as recalled by three general-purpose LLMs.

Network reactions recalled by three LLMs (Gemini 3.0, orange; GPT 5.2, blue; Claude 4.6, green) compared with a ‘Ground Truth’ literature-curated and validated fibroblast signaling network (gray reactions). LLM-generated networks used prompts based on the gene set of the Ground Truth fibroblast network. This visualization corresponds to the analyses in Figure 1D.

Figure 1—figure supplement 2
Visualization of large language model (LLM)-generated mechanosignaling networks, as recalled by three general-purpose LLMs.

Network reactions recalled by three LLMs (Gemini 3.0, orange; GPT 5.2, blue; Claude 4.6, green) compared with a ‘Ground Truth’ literature-curated and validated mechanosignaling network (gray reactions). LLM-generated networks used prompts based on the gene set of the Ground Truth mechanosignaling network. This visualization corresponds to the analyses in Figure 1E.

Figure 1—figure supplement 3
Confusion matrices, calibration analysis, and union of predicted reactions across replicates for the large language model (LLM)-generated hypertrophy networks.

(A) Confusion matrices summarizing reaction-level prediction performance for each LLM and a null predictor (hypothetical network with no connections) relative to the manually curated hypertrophy network. Actual positives were defined as reactions present in the manually curated network, whereas actual negatives were defined as possible node-to-node reactions absent from the manually curated network. (B) Calibration analyses were performed using replicate predictions for each LLM (n=10), with confidence defined as the fraction of replicates in which a reaction was predicted. (C) Venn diagram showing the overlap between the manually curated ground-truth reactions and the union of reactions predicted by at least one replicate from each LLM (confidence ≥0.1). TP, true positive; FP, false positive; TN, true negative; FN, false negative; NPV, negative predictive value; F1, F1 score.

Figure 1—figure supplement 4
Confusion matrices, calibration analysis, and union of predicted reactions across replicates for the large language model (LLM)-generated fibroblast networks.

(A) Confusion matrices summarizing reaction-level prediction performance for each LLM and a null predictor relative to the manually curated fibroblast network. Actual positives were defined as reactions present in the manually curated network, whereas actual negatives were defined as possible node-to-node reactions absent from the manually curated network. (B) Calibration analyses were performed using replicate predictions for each LLM (n=10), with confidence defined as the fraction of replicates in which a reaction was predicted. (C) Venn diagram showing the overlap between the manually curated ground-truth reactions and the union of reactions predicted by at least one replicate from each LLM (confidence ≥0.1). TP, true positive; FP, false positive; TN, true negative; FN, false negative; NPV, negative predictive value; F1, F1 score.

Figure 1—figure supplement 5
Confusion matrices, calibration analysis, and union of predicted reactions across replicates for the large language model (LLM)-generated mechanosignaling networks.

(A) Confusion matrices summarizing reaction-level prediction performance for each LLM and a null predictor relative to the manually curated mechanosignaling network. Actual positives were defined as reactions present in the manually curated network, whereas actual negatives were defined as possible node-to-node reactions absent from the manually curated network. (B) Calibration analyses were performed using replicate predictions for each LLM (n=10), with confidence defined as the fraction of replicates in which a reaction was predicted. (C) Venn diagram showing the overlap between the manually curated ground-truth reactions and the union of reactions predicted by at least one replicate from each LLM (confidence ≥0.1). TP, true positive; FP, false positive; TN, true negative; FN, false negative; NPV, negative predictive value; F1, F1 score.

Figure 2 with 1 supplement
Experimental validation of perturbation responses predicted by large language model (LLM)-generated signaling network models.

(A) Representative validations of network models generated by manual curation or by LLMs (Gemini, GPT, Claude), in comparison to experiments in conditions of Angiotensin II (AngII) or isoproterenol (ISO) from the literature (Ryall et al., 2012). (B–D) Summary of systematic validations of manually curated (Ground Truth) and LLM-generated reconstructions of hypertrophy, fibroblast, and mechanosignaling network models against perturbation experiments from the literature (n=114, 83, and 171 experiments, respectively). * Indicates p<10–11 in one-sample t-test between LLM-generated model validation scores (n=10 replicates) and ground truth model validation accuracy.

Figure 2—figure supplement 1
Experimental validation of perturbation responses predicted by full large language model (LLM)-generated signaling network models.

(A–C) Systematic validation of manually curated ground-truth models and full, unrefined LLM-generated models for the hypertrophy, fibroblast, and mechanosignaling networks against perturbation experiments from the literature (n=114, 83, and 171 experiments, respectively). Asterisks indicate p<1 × 10⁻⁷ by one-sample t-test comparing LLM-generated model validation scores across replicates (n=10 per LLM) against the corresponding ground-truth model validation accuracy. (D) Average reaction counts in the full LLM-generated network models (n=10 per network per LLM) compared with corresponding manually curated networks.

Metabolic network models generated by general-purpose large language models (LLMs) and substrate utilization analysis.

(A) Schematic of pipeline for LLM-generated models of the core E. coli metabolic network. (B) Network reactions recalled by three LLMs (Gemini, orange; GPT, blue; Claude, green) compared with a ‘Ground Truth’ manually curated metabolic network. Bar charts illustrate how frequently each reaction was predicted by the different LLMs in the replicates (n=10). Inset shows zoomed-in model coverage of the ADK1 reaction. LLM networks were generated using iterative prompts based on the GPR gene list of the core E. coli metabolic network. (C) Summary of reaction recall accuracy by Gemini, GPT, and Claude. (D) Heatmap illustrating substrate utilization predictions of the ground truth model to and LLM-generated models compared to experimental data. Color intensity within the LLM columns indicates frequency of growth prediction among the 10 replicates. (E) Summary of model performance compared to the experimental data across carbon sources. * Indicates p<10–4 in one-sample t-test between LLM-generated replicates (n=10) and the ground truth network. Within the network visualization, diamonds indicate reactions while rectangles indicate metabolites.

Tables

Author response table 1
Proportion of stimulatory and inhibitory connections present in each ground truth model and in the sets of connections predicted by each LLM (Claude, GPT, and Gemini).
Hypertrophystimulationinhibition
Ground Truth86.3913.61
Claude91.698.31
GPT89.3110.69
Gemini90.549.46

Additional files

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Jeevan Tewari
  2. Benjamin W Dahl
  3. B Adam Bates
  4. Jason A Papin
  5. Jeffrey J Saucerman
(2026)
Benchmarking biochemical networks generated by large language models
eLife 15:RP109709.
https://doi.org/10.7554/eLife.109709.3