The genetic architecture of dementia risk: how Alzheimer’s disease vulnerability converges on lipid metabolism and immune cell networks

  1. Department of Neuroscience, Karolinska Institutet, Stockholm, Sweden
  2. Science for Life Laboratory, School of Engineering Sciences in Chemistry, Biotechnology and Health, KTH Royal Institute of Technology, Stockholm, Sweden

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Hugo Bellen
    Baylor College of Medicine, Houston, United States of America
  • Senior Editor
    Ma-Li Wong
    State University of New York Upstate Medical University, Syracuse, United States of America

Reviewer #1 (Public review):

Summary:

The authors build a reusable, disease-agnostic pipeline that retrieves GWAS risk genes from the GWAS Catalog by ontology terms, filters them, and maps them onto Human Protein Atlas (HPA v24) co-expression modules at three biological scales (tissue/organ, brain region, cell type). Enrichment is assessed by a consensus of Fisher's exact test with Benjamini-Hochberg correction and 10⁶-iteration Monte Carlo simulation. Applied to AD, DLB/PD, and FTD/ALS, the analysis reports convergent neuronal-module enrichment across all three diseases, AD-specific enrichment in liver- and immune-associated modules, and DLB-specific enrichment in ciliary modules, followed by a DrugBank-based survey of compounds targeting module gene products.

Strengths:

(1) The core premise is sound and well-motivated: risk-gene lists are hard to interpret because most variants are low-penetrance and broadly expressed, and projecting them onto a multiscale expression atlas is a reasonable route from statistical association toward tissue/cell context.

(2) The pipeline is delivered as reusable, open code (GitHub) built entirely on public inputs (HPA, GWAS Catalog, DrugBank), which is a significant contribution to the field and provides opportunities for replication and expansion.

(3) The dual-enrichment design (fold enrichment with FDR correction plus a 10⁶-iteration Monte Carlo empirical null) is more defensible than any single test, and the "consensus" logic is sound.

(4) The multiscale framing (organ → brain region → cell type), with UMAP module projections, is genuinely helpful and makes the mapping legible to broad scientific backgrounds.

(5) The authors are commendably restrained on one key point: they explicitly report that most risk genes are broadly expressed and not brain-selective, rather than overstating neuronal specificity.

(6) The AD liver/immune convergence is nicely discussed and integrated, and the NAFLD-AD discussion (Kupffer-cell/hepatocyte Aβ clearance, locus coeruleus noradrenergic parallels, "type 3 diabetes") is thorough and well-referenced, even where it remains speculative.

Weaknesses:

(1) Gene-to-variant mapping via the author-reported gene field is unclear. The Methods assign genes using the GWAS Catalog author-reported gene field, which predominantly reflects the nearest gene to the lead SNP and may not be the effector gene; it is also inconsistent across studies and different time periods of publication (i.e., changing methodologies in genomics and GWAS procedures). Because every downstream result depends on the gene set, this choice likely introduces noise and bias into all enrichment, specificity, and drug claims. More modern practices link variants to genes via fine-mapping plus eQTL/pQTL colocalization, or integrative scores. At minimum, the sensitivity of the main signatures to nearest-gene versus colocalization-based assignment should be demonstrated.

(2) The suggestive threshold (p < 1×10⁻⁵) trades specificity for coverage in the analysis that requires specificity. Relaxing from 5×10⁻⁸ substantially raises the false-positive fraction of the gene set. This is defensible for exploratory coverage in under-powered DLB/FTD, but the headline claims concern disease specificity (limited gene-set overlap; distinct signatures). Non-overlap among partially false-positive lists can lead to biological specificity conclusions that may not actually exist. A genome-wide-threshold sensitivity analysis is needed to show the signatures persist. Also, see concerns below about the DLB designation.

(3) "Disease specificity" is confounded by GWAS power. AD GWAS (e.g., Bellenguez; Kunkle; Sherva ~205,500 cases) vastly outpower DLB and FTD discovery. The gene counts (453/278/219) and the minimal three-way overlap (only two genes) track sample size and locus density as much as biology. The claim of "disease-specific genetic architectures" should be tempered and ideally power-matched (e.g., subsampling AD, or restricting to comparable effective N) before specificity is asserted. Also, see concerns below about the DLB designation.

(4) Linkage disequilibrium structure at gene-dense loci is not addressed and may inflate the lipid/liver signal. The overlap genes named as driving the liver/lipid theme include APOC1, APOC2, and APOE, all of which reside within the same chromosome-19 linkage disequilibrium (along with TOMM40). Counting co-regulated, physically clustered genes from one association signal as independent risk genes risks inflation of enrichment for lipid/lipoprotein modules. The five "liver-enhanced" genes flagged in the text again lead with APOE. Evaluation of one gene per independent signal is essential before the liver/lipid signature can be interpreted as multi-gene convergence rather than a single strong gene driving the effect.

(5) The enrichment background is not clear. Risk genes are filtered to HPA brain-detected transcripts, but the Methods do not state whether the Fisher/Monte Carlo background (N_Total) is likewise restricted to brain-expressed/HPA-detected genes or reflects all protein-coding genes, which may introduce bias. Specific information on the Fisher/Monte Carlo should be provided to ensure that the enrichment background is the same. If they are not, additional analyses should be performed to ensure robustness of the findings when restricted to brain-expressed only or the broader background.

(6) Cross-scale "consensus" is not independent confirmation. The same genes reappear across tissue, brain, and cell modules, so agreement across scales is partly built-in rather than corroborating. The neuronal-signature counts (82 AD / 62 DLB / 53 FTD; 186 "unique" genes) should be accompanied by a clear statement of how much cross-scale evidence is non-redundant. Also, see concerns below about DLB designation.

(7) Temporal claims are overstated. Using control HPA tissue avoids end-stage confounds but, by construction, cannot capture disease-state programs central to AD. More importantly, framing these modules as "baseline vulnerability hotspots that precede clinical neurodegeneration" is not tested, as nothing here is longitudinal. This assumption is presented as a finding and should be reworded as a hypothesis.

(8) DLB and PD should not be merged, and the cilia-associated risk genes cannot be termed causal based on the study design. The supporting literature is almost entirely PD (Schmidt iPSC-NPCs from sporadic PD; LRRK2 PD striatum), yet DLB and PD are merged, and the signature is branded "DLB." These should not be merged, particularly as it relates to sporadic PD. Although there are similar genetic risk factors (i.e., GBA, SCNA), of which GBA is unfortunately not mentioned in the manuscript, there are major genetic differences in sporadic PD and even PD with dementia (PDD) and DLB. It is not clear why PDD was not incorporated.

Cross-sectional expression overlap with GWAS genes cannot establish that cilia-associated risk genes are "causative vulnerabilities rather than secondary effects." Please soften to association and rename to reflect the synucleinopathy grouping. Further, additional discussion of important differential genes in this category (i.e., APOE and GBA) is needed, and the pathological description of DLB is incomplete in the introduction (i.e., focuses only on synuclein).

(9) The drug/repurposing analysis rests on a weak targeting rationale. Most of the 1,777 drugs target co-expression-module neighbors of risk genes, not risk genes themselves, so "substantial repurposing reservoir" likely overstates the impact. The anesthetics example (sevoflurane/halothane/desflurane hitting GABA_A subunits and ATP2B2) is a near circular argument, as anesthetics necessarily engage neuronal ion channels. This finding does not independently confirm the neuronal signature.

Statin-dementia and hydroxychloroquine/amodiaquine links are pre-existing, and the Tirzepatide→cilia→DLB inference is highly speculative. This section should be labeled hypothesis-generating, with direct-risk-gene targets separated from module-neighbor targets.

(10) Interpretation leans heavily on nominal (p < 0.05) modules. Several of the most novel claims (parts of the liver and immune signatures) rest on nominally significant modules that do not survive FDR (itself set leniently at p_adj < 0.1). The text should make consistently explicit which claims are FDR/Monte-Carlo-supported versus nominal-only, and de-emphasize conclusions resting solely on the latter.

Reviewer #2 (Public review):

Summary:

Genes associated with risk for a specific disease commonly have widespread expression and functions across the body. Surveying patterns in these effects may reveal novel mechanisms, organs, and systems implicated in a disease, amongst other associations that are truly independent. In this work, Husen and coauthors use the Human Protein Atlas to explore such associations in Alzheimer's disease (AD), Lewy body dementia (DLB), and Frontotemporal dementia (FTD). Focusing on human non-disease tissue expression may avoid the effects of disease progression obscuring initial vulnerabilities. However, associations in non-diseased tissues do not necessarily reflect mechanisms causally related to the diseases themselves.

The work describes patterns of enrichment of genes across tissue types, brain regions, and cell types. A relatively small set of classes of each show enrichment for disease. While neural signatures are unsurprisingly prevalent, these classes are largely distinct across the three diseases. Alzheimer's disease is linked to liver and central and peripheral immune cells, while DLB shows interesting enrichments associated with cilia, which are linked to an existing literature. Results from the drug repurposing approach are then presented, with 1777 drugs linked to protein products of any of the modules enriched by the risk genes using DrugBank, categorised according to key signatures.

The authors developed R code (the HPA GeneSet Explorer) to automate the production of multi-system summaries of the organs, brain regions, cells and gene modules associated with traits and diseases and associated gene sets, within the HPA. Risk gene sets for the three dementia types were derived from the GWAS Catalog.

Strengths:

While many studies of how risk genes contribute to disease take a narrow approach focusing on organs, cell types, and processes already associated with a disease, it is a sensible approach to start with a system-agnostic approach that assesses tissues that are not ostensibly affected by disease. Here, this approach reveals a range of associations for 3 neurodegenerative diseases, identifying disease-associated modules and drug candidates that might be prioritized for subsequent confirmatory inference across biological scales. Results highlight key organs and cell types, most of which have established associations with the disease. Perhaps the most intriguing results are the links of DLB to cilia-related processes, which can be linked to some prior reports of DLB/PD but are not a core element of current theories of pathogenesis.

Weaknesses:

A difficulty with broad, multi-dataset surveys of disease associations is the need to distinguish novel and robust patterns - even if they lack causal evidence - from those that are unsurprising or do not stand out statistically. The work is exploratory in nature, but it is often hard to know how strong the evidence is for particular observations reported.

The work combines nominal, FDR<0.1, and Monte Carlo-based inference (with no apparent multiple assessment control across all tested modules) p-values throughout the paper, with patterns of effects of nominal significance. In some places, modules appear to be retained if they meet any of these criteria, muddying inference. This makes it difficult to weigh the different reported associations. Results report numbers of risk genes showing nominal p<0.05 enrichment across gene modules and biological scales - it is difficult for the reader to determine null expectations for false positives here. Similarly, it is unsurprising that thousands of drugs can be linked to the risk genes and their signatures using nominal significance.

The results have limited mechanistic specificity. The modules identified often reflect biological processes implicated in the diseases. This provides some validation of the approach, but the modules are often broadly defined, providing little mechanistic insight. For example, many aspects of ciliary biology may overlap with DLB, but can the HPA provide more specific insight? More generally, it is difficult to determine how much relevance that enrichment in non-disease tissue has for disease processes. Similarly, it is hard to determine whether overlap of drug targets from DrugBank with these modules realistically increases their prioritization.

Methodologically, there could be more detail. The paper - in particular the methods - is partially presented as a tool/pipeline paper, but thorough descriptions of the HPA models that are employed and modules reported for the analyses should still be presented in detail. The drug repurposing approach is described in a couple of sentences without a precise reference to the tool or statistical methods.

Reviewer #3 (Public review):

Summary:

The manuscript presents the development of a software tool and a computational workflow for the comparison and biological interpretation of GWAS results among three neurodegenerative diseases - AD, PD and FTD - using the data from the Human Protein Atlas. The multi-scale content analysis produces a representation of GWAS results highlighting enrichment at the level of brain regions, organ systems/tissues and cell types as well as in terms of molecular pathways. The manuscript argues that DLB, AD and FTD have differential 'modules' revealed by the procedure.

Strengths:

The system is leveraging vast knowledge sources including both the GWAS catalog and the HPA, and it is integrating these data. This synthesis of information is useful. It is making use of large investment in data generation and data warehouses in order to yield interpretation of epidemiologic results from GWAS in biological terms that may yield insight into disease processes and differences between diseases. The multiscale analysis recognizes the various biological lenses at which the implications of GWAS can be evaluated.

Weaknesses:

There is a lack of controls and/or disease comparators in the study. It is hard to assess that the workflow is performing 'as expected' without a set of positive or negative controls - or at least comparators - to gauge the performance of the tool. The statistical methods are simplistic and rely on Fisher's exact tests, UMAP analyses and clustering with little justification. There is a lack of power analysis and specification of the number of genes required for the procedure to 'work'. The mapping of GWAS hits to genes is simplistic and may in many cases be erroneous, and the implications of errors in these mappings are not considered. Many of the hits are not brain-specific - highlighting the complexity of gene function and the pleiotropic nature of gene activity. Moreover, the mapping between 'tissue enrichment' and 'tissue that is the functional driver of the GWAS signal' may be a logical flaw in the reasoning of the authors. Just because a tissue - such as liver enrichment in AD - is enriched in the GWAS gene-mapping analysis does not mean that that tissue found to be enriched is in fact the functional tissue that gave rise to the GWAS signal. Genes have different isoforms, functions, and regulatory mechanisms in different parts of the organism in different parts of development. In a phrase, the enrichment observed could be correlative and not causative, and in fact the enrichment could be driven by some hidden variable not considered. The etiologic tissue for neurodegenerative disease is the brain. The findings are not necessarily surprising or novel in the distinction between PD, AD, and FTD. Finally, the figures are perhaps not the best way to present results. There are many small pie charts that lack interesting results; the figures are in general hard to read and could use refinement in terms of fonts.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation