dbGIST: An LLM-Assisted Multi-Omics Resource for Target Exploration and Cross-Dataset Validation in Gastrointestinal Stromal Tumors

  1. Department of Gastrointestinal Surgery, The First Affiliated Hospital of Zhengzhou University, Zhengzhou, China
  2. Clinical Systems Biology Laboratories, The First Affiliated Hospital of Zhengzhou University, Zhengzhou, China
  3. Department of Pathology, The First Affiliated Hospital of Zhengzhou University, Zhengzhou, China
  4. Proteoformics Biotechnologies Inc., Zhengzhou, China
  5. School of Biological Science and Medical Engineering & School of Engineering Medicine, Beihang University, Beijing, China

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Eyal Itskovits
    GlaxoSmithKline, London, United Kingdom
  • Senior Editor
    Tony Ng
    King's College London, London, United Kingdom

Reviewer #1 (Public review):

Summary:

This tumour type is missing from the big pan-cancer databases, so none of the popular online analysis tools works for it. That's a real gap, and it's the right one to go after. The authors build an online resource that gathers the scattered public molecular datasets for this disease, adds three of their own patient cohorts, ties everything to clinical data, and exposes interactive tools, downloads, and programmatic access so other people can build on it. To show what it does, they take one gene through the whole platform - clinical, gene-expression, protein, single-cell, immune, and drug-response and then test that gene in cell lines. So there are really two things on offer here: a resource and a practical example of using it. They land very differently.

Strengths:

The resource is the real contribution, and it's done with care. It covers 37 centres and nearly 2,000 samples across five kinds of molecular data, and the authors are honest about provenance: how they screened datasets in or out, where they recorded the diagnostic codes, and why they dropped ambiguous mixed-tumour collections. The key methodological decision is the right one; every analysis runs inside its own cohort, and the cross-cohort views are explicitly "for looking, not for combining." That's exactly how you should treat heterogeneous public data, and they say so plainly instead of quietly pooling everything. Their three pathologist-confirmed cohorts add genuine independent material, so this isn't a re-skin of data that already existed. And because the code and a public access point are actually available, the reuse claim holds.

The example is internally consistent, which is what makes it persuasive. The gene reads higher in higher-risk patients across several independent cohorts and in their own protein data, tracks with the disease spreading and recurring, and lines up with worse survival. The single-cell data put it in the dividing cells; the pathway analysis points to proliferation. Three independent data types landing on the same proliferation story are the strongest part of the biology.

Weaknesses:

The honest problem is that the entire biological story rests on one gene, tested one way. The lab work is two cell lines with the gene knocked down, showing less growth and migration: there is no rescue to confirm the effect is real, no second gene to show the approach generalises, nothing in a living animal. That earns the modest claim: the resource can point you at a candidate worth testing. It does not earn the headline claim that the platform reliably generates good target hypotheses, because we only ever watch it succeed once. One example illustrates a workflow; it doesn't establish a method.

Some of the statistics won't survive scrutiny. The clearest case is a perfect separation between treatment-resistant and treatment-sensitive cases from a single immune cell population, reported with no error bars, no check for information leakage, and apparently from very few samples. A perfect result in that setting is almost always overfitting or a small-sample artefact, not a strong classifier. The same pattern shows up elsewhere: small groups, p-values with no effect sizes or error bars, and no correction for the enormous number of features and cohorts being tested across the whole platform. Separately, one drug result is a correlation against a predicted sensitivity score from a model.

The AI assistant gets far more weight than the evidence supports. Credit where due: the authors are clear and consistent that it only helps interpret and navigate, and never touches the data, the statistics, or the results. That's the correct line to draw, and they hold it. But the assistant itself is never tested, no accuracy numbers, no benchmark, no error analysis, no described way for a human to check what it produces. Calling it something that "fundamentally transforms the user experience" is an assertion, not a finding. And since even the literature feature is admitted not to be a proper systematic review, the prominence of the artificial-intelligence framing runs ahead of what's been shown.

Reviewer #2 (Public review):

Summary:

dbGIST appears to be the first dedicated multi-omics resource worldwide that is specifically focused on GIST.

Strengths:

The main value of the paper is not simply that the authors collected datasets, but that they built a usable resource around them, with cohort-aware analyses, curated clinical labels, interactive visualizations, downloadable results, selected API access, and an optional LLM-assisted interface. The work is solid, and the database is likely to be useful for GIST researchers interested in target discovery, cross-dataset validation, drug-response hypotheses, and translational follow-up.

The MCM7 analysis is a reasonable use case. It shows how a user can start from one candidate gene and then move across transcriptomic, proteomic, clinical, single-cell, immune-related, drug-response, and experimental evidence. I do not see this as the main discovery of the paper, but rather as a practical demonstration of what the database can do. That is appropriate for a resource manuscript.

Weaknesses:

(1) The authors should make the organization of the platform a little easier to follow. The manuscript refers to five primary omics layers, six omics-focused pages, and eight analytical modules. This structure is understandable after reading the relevant sections, but it may not be immediately obvious to readers. A brief clarification of how the omics layers, web pages, and analytical modules relate to each other would help.

(2) Since dbGIST is a live web resource, the authors should provide a clear versioning statement. The manuscript should indicate which version of the database corresponds to the analyses and figures reported in the paper, and how future updates will be distinguished from the version evaluated here. This is a small point, but it matters for reproducibility.

(3) The API function is a strength of the resource, but it is still described rather generally. The authors should give more concrete documentation of what can be accessed through the API, what inputs are required, and what type of output is returned. This could be placed in the supplementary materials. It would make the database more useful for computational users.

(4) The manuscript should clarify the status of downloadable data. It is clear that figures, source-data tables, and selected derived outputs are available, but it is less clear whether the full processed matrices used internally by the platform are downloadable or only maintained for deployment. This distinction should be stated plainly.

(5) The statistical reporting in the MCM7 clinical-association analyses needs a little more care. Several p-values are shown across different cohorts and clinical variables. The authors should state whether these are nominal p-values or adjusted p-values. If they are nominal, that is acceptable for a resource demonstration, but the exploratory nature of the analyses should be made clear.

(6) The ROC analyses for imatinib response should include sample sizes, and confidence intervals for AUC values would be useful if available. Some of the AUC values are high, and without group sizes, it is difficult to judge how stable those estimates are. The authors should avoid implying that these ROC results are validated predictive models.

(7) The interpretation of MCM7 should be slightly more cautious. MCM7 is a well-known DNA replication and cell-cycle gene, and the single-cell analyses seem to support its association with proliferative cell states. This is biologically consistent, but it also means that MCM7 expression should not be presented as tumour-cell-specific without qualification. The manuscript should frame it mainly as a proliferation-associated signal in the current analysis.

(8) The drug-response section would benefit from a clearer explanation of the response metric. The authors report correlations between MCM7 expression and predicted response to C6-ceramide, but readers need to know whether the predicted value represents IC50, AUC, sensitivity score, or another metric. The direction of interpretation should also be made explicit, since a negative correlation can mean different things depending on the scoring system.

(9) The single-cell annotation would be more convincing if the authors provided a compact marker-gene summary for the major cell types in each single-cell cohort. The current description of annotation by source labels, marker inspection, and manual curation is reasonable, but users of the database would benefit from seeing the marker evidence behind the labels.

(10) The experimental validation section should include a few routine details that are currently not easy to find. The siRNA sequences or target regions, number of biological replicates, statistical tests for the CCK-8 and wound-healing assays, and details of wound-closure quantification should be reported. These additions would make the in vitro part more reproducible.

(11) The wound-healing result should be interpreted with caution. Since MCM7 knockdown reduces proliferation, reduced wound closure could reflect changes in proliferation, migration, or both. Unless proliferation was controlled during the wound-healing assay, the authors should avoid describing this result as purely migratory.

(12) The LLM-related claims should remain conservative. The assistant is a useful feature for navigation, plain-language explanation, and user support, especially for clinicians or wet-lab researchers. However, the strongest statements about the LLM transforming interpretation or automating analysis should be toned down. The important point is that the LLM layer helps users interact with the resource, while the numerical analyses come from predefined dbGIST modules.

Reviewer #3 (Public review):

Summary:

The dbGist dataset/tool would provide substantial value to the cancer research community.

Strengths:

The manuscript presents dbGIST, a dedicated GIST-focused multiomics resource integrating data from 37 centers and ~2k samples across genomics, transcriptomics, proteomics, phosphoproteomics, and single-cell transcriptomics. Given that GIST is virtually absent from major cancer genomics consortia (TCGA, ICGC), this resource fills a genuine gap and represents a valuable contribution to the GIST research community.

(1) The MCM7 case study effectively demonstrates the platform's utility, linking a resource-derived candidate to survival outcomes.

(2) The LLM-assisted interface (dbGIST Assistant) is a reasonable addition for accessibility, lowering the barrier for clinicians and wet-lab researchers, who may not always have the skill set required for proper data analysis, especially for a rich and wide dataset like the dataset in question.

Weaknesses:

(1) Data deposition (major):

While the manuscript references public accessions for raw source datasets and provides a GitHub repository for code, it remains unclear where the **curated, harmonized data matrices** - which represent the core value-add of this work - are independently deposited. Access to these processed data appears to depend entirely on the dbGIST web interface and API. The authors should deposit the harmonized matrices in a persistent, general-purpose repository to ensure long-term availability independent of the web platform.

(2) LLM agent capabilities underspecified:

The manuscript would benefit from a clearer description of the assistant's capabilities and boundaries. Specifically, what tools or actions are available to the LLM agent? Can it execute code against the underlying data, trigger analytical modules programmatically, or is it limited to natural-language explanation of pre-computed results? Clarifying this would help readers assess the scope of the AI layer and distinguish it from agentic platforms that perform computation on behalf of the user.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation