Population-scale proteome variation in human induced pluripotent stem cells

  1. Bogdan Andrei Mirauta
  2. Daniel D Seaton
  3. Dalila Bensaddek
  4. Alejandro Brenes Murillo
  5. Marc Jan Bonder
  6. Helena Kilpinen
  7. HipSci Consortium
  8. Oliver Stegle  Is a corresponding author
  9. Angus I Lamond  Is a corresponding author
  1. European Bioinformatics Institute, United Kingdom
  2. University of Dundee, United Kingdom
  3. University College London, United Kingdom
  4. European Molecular Biology Laboratory, European Bioinformatics Institute, United Kingdom

Abstract

Human disease phenotypes are ultimately driven primarily by alterations in protein expression and/or function. To date, relatively little is known about the variability of the human proteome in populations and how this relates to variability in mRNA expression and to disease loci. Here, we present the first comprehensive proteomic analysis of human induced pluripotent stem cells (iPSC), a key cell type for disease modelling, analysing 202 iPSC lines derived from 151 donors, with integrated transcriptome and genomic sequence data from the same lines. We characterised the major genetic and non-genetic determinants of proteome variation across iPSC lines and assessed key regulatory mechanisms affecting variation in protein abundance. We identified 654 protein quantitative trait loci (pQTLs) in iPSCs, including disease-linked variants in protein coding sequences and variants with trans regulatory effects. These include pQTL linked to GWAS variants that cannot be detected at the mRNA level, highlighting the utility of dissecting pQTL at peptide level resolution.

Data availability

RNA-Seq data for 331 samples are available on the European Nucleotide Archive (ENA): study PRJEB7388; accession ERP007111. Proteomics quantifications (protein group and peptide resolution; MaxQuant output), and run parameters are available on the PRIDE Archive PRIDE (PXD010557). Analysed data is included in the supplementary external files.

The following data sets were generated

Article and author information

Author details

  1. Bogdan Andrei Mirauta

    Statistical genomics, European Bioinformatics Institute, Cambridge, United Kingdom
    Competing interests
    The authors declare that no competing interests exist.
  2. Daniel D Seaton

    Statistical genomics, European Bioinformatics Institute, Cambridge, United Kingdom
    Competing interests
    The authors declare that no competing interests exist.
  3. Dalila Bensaddek

    Centre for Gene Regulation & Expression, School of Life Sciences, University of Dundee, Dundee, United Kingdom
    Competing interests
    The authors declare that no competing interests exist.
  4. Alejandro Brenes Murillo

    Centre for Gene Regulation & Expression, School of Life Sciences, University of Dundee, Dundee, United Kingdom
    Competing interests
    The authors declare that no competing interests exist.
  5. Marc Jan Bonder

    Statistical genomics, European Bioinformatics Institute, Cambridge, United Kingdom
    Competing interests
    The authors declare that no competing interests exist.
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0002-8431-3180
  6. Helena Kilpinen

    Great Ormond Street Institute of Child Health, University College London, London, United Kingdom
    Competing interests
    The authors declare that no competing interests exist.
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0001-6692-6154
  7. HipSci Consortium

  8. Oliver Stegle

    Wellcome Trust Genome Campus, European Molecular Biology Laboratory, European Bioinformatics Institute, Cambridge, United Kingdom
    For correspondence
    oliver.stegle@ebi.ac.uk
    Competing interests
    The authors declare that no competing interests exist.
  9. Angus I Lamond

    Centre for Gene Regulation and Expression, University of Dundee, Dundee, United Kingdom
    For correspondence
    a.i.lamond@dundee.ac.uk
    Competing interests
    The authors declare that no competing interests exist.
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0001-6204-6045

Funding

Wellcome Trust Strategic Award and UK Medical Research Council (WT098503)

  • Bogdan Andrei Mirauta
  • Daniel D Seaton
  • Dalila Bensaddek

Wellcome Trust Strategic Award (105024/Z/14/Z)

  • Bogdan Andrei Mirauta

The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication.

Copyright

© 2020, Mirauta et al.

This article is distributed under the terms of the Creative Commons Attribution License permitting unrestricted use and redistribution provided that the original author and source are credited.

Metrics

  • 4,078
    views
  • 450
    downloads
  • 48
    citations

Views, downloads and citations are aggregated across all versions of this paper published by eLife.

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Bogdan Andrei Mirauta
  2. Daniel D Seaton
  3. Dalila Bensaddek
  4. Alejandro Brenes Murillo
  5. Marc Jan Bonder
  6. Helena Kilpinen
  7. HipSci Consortium
  8. Oliver Stegle
  9. Angus I Lamond
(2020)
Population-scale proteome variation in human induced pluripotent stem cells
eLife 9:e57390.
https://doi.org/10.7554/eLife.57390

Share this article

https://doi.org/10.7554/eLife.57390

Further reading

    1. Genetics and Genomics
    Jongkeun Park, WonJong Choi ... Dongwan Hong
    Research Article

    An unprecedented amount of SARS-CoV-2 data has been accumulated compared with previous infectious diseases, enabling insights into its evolutionary process and more thorough analyses. This study investigates SARS-CoV-2 features as it evolved to evaluate its infectivity. We examined viral sequences and identified the polarity of amino acids in the receptor binding motif (RBM) region. We detected an increased frequency of amino acid substitutions to lysine (K) and arginine (R) in variants of concern (VOCs). As the virus evolved to Omicron, commonly occurring mutations became fixed components of the new viral sequence. Furthermore, at specific positions of VOCs, only one type of amino acid substitution and a notable absence of mutations at D467 were detected. We found that the binding affinity of SARS-CoV-2 lineages to the ACE2 receptor was impacted by amino acid substitutions. Based on our discoveries, we developed APESS, an evaluation model evaluating infectivity from biochemical and mutational properties. In silico evaluation using real-world sequences and in vitro viral entry assays validated the accuracy of APESS and our discoveries. Using Machine Learning, we predicted mutations that had the potential to become more prominent. We created AIVE, a web-based system, accessible at https://ai-ve.org to provide infectivity measurements of mutations entered by users. Ultimately, we established a clear link between specific viral properties and increased infectivity, enhancing our understanding of SARS-CoV-2 and enabling more accurate predictions of the virus.

    1. Cell Biology
    2. Genetics and Genomics
    Showkat Ahmad Dar, Sulochan Malla ... Manolis Maragkakis
    Research Article

    Cells react to stress by triggering response pathways, leading to extensive alterations in the transcriptome to restore cellular homeostasis. The role of RNA metabolism in shaping the cellular response to stress is vital, yet the global changes in RNA stability under these conditions remain unclear. In this work, we employ direct RNA sequencing with nanopores, enhanced by 5ʹ end adapter ligation, to comprehensively interrogate the human transcriptome at single-molecule and -nucleotide resolution. By developing a statistical framework to identify robust RNA length variations in nanopore data, we find that cellular stress induces prevalent 5ʹ end RNA decay that is coupled to translation and ribosome occupancy. Unlike typical RNA decay models in normal conditions, we show that stress-induced RNA decay is dependent on XRN1 but does not depend on deadenylation or decapping. We observed that RNAs undergoing decay are predominantly enriched in the stress granule transcriptome while inhibition of stress granule formation via genetic ablation of G3BP1 and G3BP2 rescues RNA length. Our findings reveal RNA decay as a key component of RNA metabolism upon cellular stress that is dependent on stress granule formation.