Imputation of structural variants using a multi-ancestry long-read sequencing panel enables identification of disease associations

  1. Global Computational Biology and Digital Sciences (gCBDS), Boehringer Ingelheim Pharma GmbH & Co. KG, Biberach an der Riss, Germany
  2. Institute of Cancer and Genomic Sciences, University of Birmingham, Birmingham, United Kingdom
  3. BI X GmbH, Ingelheim am Rhein, Germany
  4. Department of Neurology, School of Medicine, Technical University of Munich, Munich, Germany
  5. Gencove, New York, United States
  6. Discovery Research, Boehringer Ingelheim Pharma GmbH & Co. KG, Ingelheim am Rhein, Germany

Peer review process

Revised: This Reviewed Preprint has been revised by the authors in response to the previous round of peer review; the eLife assessment and the public reviews have been updated where necessary by the editors and peer reviewers.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Jun Kim
    Chungnam National University, Daejeon, Republic of Korea
  • Senior Editor
    Murim Choi
    Seoul National University, Seoul, Republic of Korea

Reviewer #1 (Public review):

Summary:

The authors sequenced 888 individuals from the 1000 Genomes Project using the Oxford Nanopore long-read sequencing method to achieve highly sensitive, genome-wide detection of structural variants (SVs) at the population level. They conducted solid benchmarking of SV calling and systematically characterized the identified SVs. While short-read sequencing methods, including those used in the 1000 Genomes Project, have been widely applied, they exhibit high accuracy in detecting single nucleotide variants (SNVs) and small insertions and deletions but have limited sensitivity for SV detection. This study significantly enhances SV detection capabilities, establishing it as a valuable resource for human genetic research. Furthermore, the authors constructed an SV imputation panel using the generated data and imputed SVs in 488,130 individuals from the UK Biobank. They then conducted a proof-of-principle genome-wide association study (GWAS) analysis based on the imputed SVs and selected traits within the UK Biobank. Their findings demonstrate that incorporating SV-GWAS analysis provides additional insights beyond conventional GWAS frameworks focusing on SNVs, particularly in improving fine-mapping.

Strengths:

The authors constructed a high-sensitivity reference panel of genome-wide SVs at the population level, addressing a critical gap in the field of human genetics. This resource is expected to significantly advance research in human genetics. They demonstrated the imputation of SVs in individuals from the UK Biobank using this panel and conducted a proof-of-concept SV-based GWAS. Their findings highlight a novel and effective strategy for integrating SVs into GWAS, which will facilitate the analysis of human genetic data from the UK Biobank and other datasets. Their conclusions are supported by comprehensive analyses.

Weaknesses:

The authors have addressed many of my previous comments, and I appreciate their efforts. However, I still have two related concerns.

(1) Shortly after reviewing this manuscript last year, my laboratory obtained access to the UK Biobank (UKB) Tier 3 dataset for an unrelated project. In August 2025, I searched the UKB Research Analysis Platform (UKB-RAP) for the imputed structural variant (SV) dataset described in this manuscript but was unable to locate it. After contacting UKB, I was informed that they were developing the system for releasing the data. To the best of my knowledge, the dataset remains unavailable. A major contribution of this work is the generation of an imputed SV resource for approximately 500,000 UKB participants with extensive phenotypic information. If this resource is not accessible to the research community, even to authorized UKB users, the practical impact and utility of the study are substantially diminished.

(2) Given that the imputed SV dataset is currently unavailable, it becomes even more important for the authors to provide a detailed, ready-to-run SV imputation pipeline for UKB-RAP, even if the "data processing simply consisted of running standard bioinformatics tools with the parameters exactly as described in the manuscript". In particular, the pipeline should include practical information such as computational requirements (e.g., memory and storage), expected running time, and estimated cost. Anyone with experience using UKB-RAP will agree that reproducing large-scale analyses on the platform can be both technically complex and financially expensive. Such pipeline would greatly improve the reproducibility and accessibility of this work.

Because my initial assessment of the manuscript was generally positive, I do not wish to change my overall evaluation, summary, or assessment of its strengths. However, I would view the work even more favorably if either (i) the imputed SV dataset became publicly available to authorized UKB users, or (ii) the authors extended their SV-GWAS analyses to the full range of UKB phenotypes and released the resulting summary statistics, analogous to the Pan UKBB ("https://pan.ukbb.broadinstitute.org/") resource. Although this would require considerable additional effort and computational resources, it would substantially enhance the long-term value and impact of the study.

Finally, I would like to emphasize that these comments are not intended to create unnecessary difficulties for the authors or the editors. Rather, I believe this highlights a broader issue in the use of this kind of large public datasets: reviewers cannot independently verify key results, and readers cannot readily build upon the work if the underlying resources are inaccessible, even after obtaining authorized access to the original dataset. I hope the authors, together with the eLife editors and UK Biobank where appropriate, can help facilitate the timely release of this valuable resource.

Author response:

The following is the authors’ response to the original reviews.

Our revision includes:

(1) The generated data from long-read whole-genome sequencing of 1000 Genomes Project samples, including FASTQ files, SV calls, and the imputation panel, are now openly available via ENA and OpnMe. The imputed structural variant data have been submitted to UK Biobank for release through the UK Biobank Research Analysis Platform, subject to UK Biobank release procedures. SV-WAS summary statistics have been made available via OpnMe.

(2) Clarification of analyses and methods, addition of two new Supplementary Figures, and correction of minor issues throughout the manuscript.

(3) A significantly expanded Discussion to address the reviewers’ comments and better contextualise our methods and results.

eLife Assessment

This fundamental work significantly enhances our understanding of how structural variants influence human phenotypes. The conclusion is convincingly supported by rigorous analyses of long-read sequencing data. If the raw data are made publicly available, these high-quality datasets and findings will further advance our knowledge of genetic variation in the human population.

We thank the editors for this positive assessment of our work. The raw long-read sequencing data (FASTQ files) can now be accessed through the European Nucleotide Archive (ENA) under accession number PRJEB89727, as part of a larger collection of 1019 sequenced probands from the 1000 Genomes Project (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727). The generated imputation panel and the structural variant calls, based on the 888 probands used in the present manuscript, remain freely available for download at https://opnme.com/genomiclens. We have now added the summary statistics of 32 SV-wide association studies to the same resource. In addition, we have submitted the imputed SV genotypes for UK Biobank participants to the UK Biobank; once processed by UK Biobank, these genotypes will be released via the UK Biobank Research Analysis Platform (RAP).

Public Reviews:

Reviewer #1 (Public review):

Summary:

The authors sequenced 888 individuals from the 1000 Genomes Project using the Oxford Nanopore long-read sequencing method to achieve highly sensitive, genome-wide detection of structural variants (SVs) at the population level. They conducted solid benchmarking of SV calling and systematically characterized the identified SVs. While short-read sequencing methods, including those used in the 1000 Genomes Project, have been widely applied, they exhibit high accuracy in detecting single nucleotide variants (SNVs) and small insertions and deletions but have limited sensitivity for SV detection. This study significantly enhances SV detection capabilities, establishing it as a valuable resource for human genetic research. Furthermore, the authors constructed an SV imputation panel using the generated data and imputed SVs in 488,130 individuals from the UK Biobank. They then conducted a proof-of-principle genome-wide association study (GWAS) analysis based on the imputed SVs and selected traits within the UK Biobank. Their findings demonstrate that incorporating SV-GWAS analysis provides additional insights beyond conventional GWAS frameworks focusing on SNVs, particularly in improving fine mapping.

The authors constructed a high-sensitivity reference panel of genome-wide SVs at the population level, addressing a critical gap in the field of human genetics. This resource is expected to significantly advance research in human genetics. They demonstrated the imputation of SVs in individuals from the UK Biobank using this panel and conducted a proof-of-concept SV-based GWAS. Their findings highlight a novel and effective strategy for integrating SVs into GWAS, which will facilitate the analysis of human genetic data from the UK Biobank and other datasets. Their conclusions are supported by comprehensive analyses.

We thank the reviewer for highlighting the value of our SV imputation reference panel.

Weaknesses:

(1) Although the authors employ state-of-the-art analytical approaches for the identification of SVs, the overall accuracy remains suboptimal, as indicated by an F1 score of 74.0%, particularly in tandem repeat regions. To enhance accuracy, it would be beneficial to explore alternative SV detection methods or develop novel approaches. Given the value of the reference panel and the fact that improved SV accuracy would lead to more precise SV imputation and GWAS results, investing effort in methodological refinement is highly encouraged.

Accurate SV calling remains an active area of research and is beyond the scope of the present study. Tandem repeat regions are particularly challenging for standardised SV detection. We believe that achieving a benchmark for NA12878 of F1 = 74% on a genome-wide level and, notably, F1 = 91% when excluding longer tandem repeats, represents strong performance. This result is especially convincing when considering that our benchmarking compared the SV calls to data generated using a different sequencing technology and processed using different bioinformatics pipelines.

(2) From the Methods section, it appears that the authors employed Beagle for both imputation and the UK Biobank imputation.

(a) It would be better to explicitly clarify this in the Results section and provide a detailed description of the corresponding procedures and parameters in the Methods section for both analyses, as this represents a key aspect of the study.

We thank the reviewer for these suggestions. Accordingly, we added the clarification to the Methods section that the leave-one-out imputation used exactly the same pipeline and settings as the UK Biobank imputation (page 14, section “Leave-one-out imputation performance”):

“We excluded one individual from the panel and imputed SVs for this individual using the panel of the remaining 887 samples, applying exactly the same pipeline and settings as those later used for SV imputation into UK Biobank (see below).”

(b) Additionally, Beagle is not specifically designed for SV imputation, the imputation quality of SVs is generally lower than that of SNVs. Exploring strategies to improve SV imputation, such as developing a novel method with reference panel data, may enhance performance.

As stated in our manuscript (page 4), we believe that, in our study, the imputation quality of SVs is lower than of that of SNVs primarily because of a) the greater difficulty of SV calling compared to SNV genotype calling and b) the heterogeneity in SV representation across samples. Once SVs are encoded as bi-allelic markers in the reference panel, they can be imputed using the same LD/haplotype-based framework as any other variants. Accordingly, improving SV imputation is likely to benefit most from more accurate upstream SV calling and genotyping (e.g., through more robust multi-sample calling and harmonised variant representations) and not so much from improved or SV-specific imputation methods. While improved imputation is an important research direction, it is beyond the scope of the present manuscript.

(c) It is also important to assess how this reduced imputation quality may influence GWAS results. For instance, it would be useful to examine whether associated SVs exhibit higher imputation quality and whether SVs with lower quality are less likely to achieve significant association signals. In addition, the lower imputation quality observed for INV, DUP, and BND variants (Figure 3) may be due to their greater lengths (Figure 2). It is better to investigate the relationship between SV length and imputation quality.

We agree that imputation quality can influence GWAS results. For example, for the FEV1/FVC phenotype, SVs with INFO > 0.9 are almost twice as likely to reach genome-wide significance (p < 5e-8) compared with SVs with 0.7 < INFO < 0.9 (odds ratio 1.95; Fisher’s exact test p-value 2.5e-5). This is consistent with the intuitive notion (applicable to any variants, not only to SVs) that greater uncertainty in the imputed genotypes dilutes association signals and therefore reduces power. For a detailed discussion of the relationship between allele frequency, imputation accuracy, and GWAS association results, see Zhang et al., Human Molecular Genetics 31(1):146–155 (2022), https://doi.org/10.1093/hmg/ddab203

We have now investigated the relationship between SV length and imputation quality (the new Supplementary Figure 6). The results suggest that the observed association between imputation quality and SV size is primarily driven by the SV-size–dependent minor allele frequency in the imputation panel.

(3) All examples presented in the manuscript focus on SVs that overlap with genes. It may also be valuable to investigate SVs that do not overlap with genes but intersect with enhancer regions. SVs can contribute to disease by altering regulatory elements, such as enhancers, which play a crucial role in gene expression. Including such analyses would further demonstrate the utility of SV-GWAS and provide deeper insights into the functional impact of SVs.

We agree with the reviewer that examining SVs intersecting with enhancer regions could be an interesting direction for future studies, as it would provide additional insights into regulatory mechanisms and disease associations. However, in the present proof-of-principle study, we prefer focusing on SVs overlapping with genes and have now highlighted this in additional detail in the revised manuscript (Discussion, page 7):

“In the present proof-of-principle study, we focused on SVs overlapping with the coding sequence of genes. In future applications of our SV imputation panel, more refined gene mapping approaches could be employed, e.g., including SVs overlapping enhancer regions or epigenetic marks. Such an enhanced mapping would increase the number of identified associated genes and thus provide additional insights into regulatory mechanisms and disease biology.”

(4) The data availability link currently provides only a VCF file ("sniffles2_joint_sv_calls.vcf.gz") containing the identified SVs.

(a) It would be beneficial for the authors to make all raw sequencing data (FASTQ files) and key processed datasets (such as alignment results and merged SV and SNV files) available. Providing these resources would enable other researchers to develop improved SV detection and imputation methods or conduct further genetic analyses.

Thank you for emphasising the importance of data sharing, which we agree with.

The Data Availability section of the manuscript already includes a link to https://opnme.com/genomiclens, where we made both the SV calls and the full and reduced SV imputation panel files freely available. We have now added SV summary statistics from 32 SVwide association studies to the same resource. Based on the reviewer’s request, we now also reference the ENA repository project PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727), where the raw FASTQ files are available for download, in the manuscript.

We have appended the Data Availability statement on page 22 of the revised manuscript as follows:

“Raw SV calls, the long-read sequencing-based SV imputation panel, and the SV summary statistics from 32 SV-wide association studies are available through the OpnMe initiative of Boehringer Ingelheim GmbH (https://opnme.com/genomiclens). The raw long-read sequencing data (FASTQ files) for the 1000 Genomes Project samples included in this study are accessible via the European Nucleotide Archive under accession number PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727). The dataset analysed here constitutes a subset of this broader collection.”

(b) Furthermore, establishing a dedicated website for data access, along with a genome browser for SV visualization, could significantly enhance the impact and accessibility of the study. Additionally, all code, particularly the SV imputation pipeline accompanied by a detailed tutorial, should be deposited in a public repository such as GitHub. This would support researchers in imputing SVs and conducting SV-GWAS on their own datasets.

The Methods section provides a full and detailed description of the imputation pipeline and parameters in the section “Preprocessing and imputation of SVs into UK Biobank” on page 14 of the revised manuscript. Our data processing simply consisted of running standard bioinformatics tools with the parameters exactly as described in the manuscript.

Reviewer #1 (Recommendations for the authors):

(1) In the Results section, Figure 3b is mentioned before 3a, and it is better to switch them in the Figure.

Thank you for highlighting this fact. We acknowledge that typically the sequence of sections matches exactly between text and figures. However, in this specific case, we would prefer to deviate from the norm: In our opinion, Figure 3 is easier to interpret in its current sequence. At the same time, the text flows more logically in its current sequence, describing 3b before 3a. We would therefore prefer to stick to the current order, even if it means that Fig. 3b is described before 3a in the text.

(2) Page 10, "Figure 1e" -> "Figure 2e".

Thank you, we corrected this issue.

(3) Page 14, "Leave-one out" -> "Leave-one-out".

Thank you, we corrected this mistake.

(4) It is better not to use abbreviations in the subheadings, especially "UKB" (page 3).

Thank you, we changed the acronym ‘UKB’ to ‘UK Biobank’ in all subheadings.

Reviewer #2 (Public review):

Summary:

The authors aimed to develop a novel and efficient method for SV detection, utilizing data from the 1000 Genomes Project (1KGP) for modeling and calibration. This method was subsequently validated using UK population data and applied to identify structural variants associated with specific disease phenotypes.

Strengths:

Third-generation single-molecule sequencing data offers several advantages over traditional high-throughput sequencing methods, particularly due to its long-read lengths, which provide valuable insights into significant forms of genomic variation. The authors have developed an efficient method for detecting structural variations and optimizing the utilization of genomic data. We hope that this method will continue to be refined, enabling researchers to more effectively leverage long-read data, high-throughput data, or even a synergistic combination of both.

Weaknesses:

Although this research contributes to our ability to more effectively utilize long-length and high-throughput data, there are some key issues that need to be addressed in terms of analyzing the specific results as well as writing the article.

Reviewer #2 (Recommendations for the authors):

(1) How to discuss the lower detection rate of structural variations (SVs) in East Asian populations, it is worth considering whether the authors' training dataset, which may have been based on raw data with insufficient representation of East Asian individuals, could have introduced a bias favoring other populations. This potential bias might arise from the relatively limited data available for Asian ancestry. Alternatively, the observed differences could also be influenced by the role of natural selection, which may have shaped the genomic landscape of East Asian populations in distinct ways. Further investigation is needed to clarify these possibilities.

Thank you for raising this important point. Although an interesting research direction, a detailed investigation of the factors affecting SV detection rates is beyond the scope of the present study. However, we do not think that the lower detection rate in East Asians is due to an underrepresentation of Asian ancestry in our dataset. To explain this to all readers, we have added the following explanation to page 7 of the Discussion:

“In this context, we observed that the number of SVs detected per individual differed between superpopulations. We identified the highest average number of SVs in individuals of African descent and a slightly lower average in East Asians, compared to the other superpopulations. While we included a higher number of African ancestry individuals, the number of East Asian individuals included in our reference panel was comparable to the number of individuals from other, non-African ancestries. In fact, it was even larger than the number of European ancestry individuals (AFR n=241, SAS n=171; EAS n=168; EUR n=164; AMR n=144). Therefore, we do not expect a major bias from underrepresentation of any superpopulation in the training dataset. It is well established that African ancestry is more diverse than is the case for other superpopulations [32, 33] and previous studies indicate that East Asian populations tend to exhibit slightly lower genetic diversity compared to European populations [34], which is consistent with the lower observed SV counts per genome.”

(2) The authors did not present the results of the detection of CNV.

Copy number variations (CNVs) are considered a subclass of structural variants. In our analysis, we detected deletions and duplications, which represent the most common forms of CNVs. However, we did not specifically investigate high copy-number SVs, as these are often larger than what can be reliably detected using long-read sequencing. Large-scale CNVs are typically identified in biobank studies through analysis of intensity data from genotyping microarrays using tools like PennCNV, and there is extensive literature supporting the use of this microarray approach in UK Biobank and other genotyped cohorts, see for example Aguirre et al.: Phenomewide Burden of Copy-Number Variation in the UK Biobank. Am J Hum Genet. 2019, Aug 1;105(2):373-383. doi:10.1016/j.ajhg.2019.07.001.

(3) Multiple testing correction is essential for ensuring the validity of large-scale structural variation (SV) association analyses. It is strongly recommended that the statistical methods and correction strategies employed, such as Bonferroni correction or false discovery rate (FDR) control, be explicitly detailed to enhance the transparency and reliability of the findings.

For genome-wide SV association analyses, we applied the commonly used genome-wide significance threshold of 5e-8, which is standard in genome-wide studies. Given that these were exploratory proof-of-principle analyses illustrating use cases for SV analyses, we decided not to correct on top of that for multiple testing for the number of traits (32) tested. For the pQTL analyses, we further adjusted this threshold using a Bonferroni-type correction based on the number of proteins tested (1,463), to account for the increased number of multiple comparisons.

We have now added a more detailed description of this multiple testing procedure to the Methods subsection “SV-wide association studies in UK Biobank” on page 16 of the revised manuscript:

“In the exploratory SV-WAS, we used the standard threshold for genome-wide significance of p < 5×10-8. For the pQTL analyses, we applied Bonferroni correction for multiple testing on top of that genome-wide threshold, correcting for the number of tested protein levels (n=1463): p < 5×10-8/ 1463 = 3.4×10-11.”

(4) The study primarily relied on data from the 1000 Genomes Project (1KGP) and the UK Biobank; however, the UK Biobank cohort is predominantly composed of individuals of European ancestry, which may restrict the generalizability of the research findings to other populations.

Our reference panel was constructed to cover multiple ancestries, enabling imputation for diverse populations. Thus, our imputation panel can be applied to biobanks around the world and is freely available for this purpose. As a proof of principle, we have demonstrated the feasibility and performance of SV imputation in UK Biobank as an example of a broadly accessible cohort. We are looking forward to biobanks from diverse ancestries downloading our imputation panel and applying it to their populations.

(5) Although the study employed long-read sequencing technology, the validation of structural variation (SV) detection accuracy predominantly relied on internal data, such as 'leave-one-out' validation. To further strengthen the reliability of the SV detection methods, it is recommended to incorporate additional external independent datasets for validation.

The leave-one-out procedure in our study was used to validate the imputation performance, not the accuracy of SV detection. To assess SV calling accuracy, we performed extensive benchmarking against external SV call datasets derived from PacBio long-read sequencing and Illumina short-read sequencing. These details are provided under the subheading ‘Structural variant calling and benchmarking’ in the Results section on page 2 of the manuscript.

(6) Some of the SVs mentioned in the study overlap with disease association loci in the GWAS Catalog, but functional annotation and exploration of the biological mechanisms of these SVs are more limited. It is suggested that LD can be added to analyse whether there are SNP that are highly linked to them to further explore their functions.

We thank the reviewer for this suggestion. We have actually conducted an analysis addressing exactly this question: We performed conditional association analyses of the SV signals with nearby short variants (SNPs and InDels) at the SV locus. Such a conditional analysis addresses whether the observed SV association is influenced by LD-correlated SNPs or not. The results of this analysis are reported in Supplementary Tables 16 and 17. These tables include both the conditional analysis results and the LD between each SV and the variant at the locus with the second-highest evidence for an association.

Researchers interested in exploring the biological significance of the SV-WAS results in more detail can now download the full SV-WAS summary statistics from https://opnme.com/genomiclens.

(7) The discussion section could be further expanded to explore the role of SV in complex diseases and its potential application in precision medicine. For example, it could discuss how SV information can be integrated into existing GWAS frameworks to enhance the accuracy of disease risk prediction.

Thank you for the suggestion, we have now added the following sentences to the discussion (page 7/8):

“Structural variants can influence complex disease biology through either the disruption of coding sequence or an altered regulation of gene expression. Such effects may not be well captured by short variants alone. Incorporating SVs into GWAS and follow-up analyses would thus provide more accurate disease risk prediction, uncover underlying pathomechanisms by highlighting actionable pathways and targets, and support precision medicine by providing biomarkers for patient stratification.”

(8) The geographic labeling of certain samples in Figure 2 appears to contain inaccuracies. For instance, the CDX sample, which represents the Dai population from Xishuangbanna in China's Yunnan Province, is currently mislabeled as originating from China's Inner Mongolia. This discrepancy should be corrected to ensure the accuracy of the data representation.

We apologise for the misunderstanding. The geographic map in Figure 2a serves as an illustrative mapping of the samples to countries. It is intended to provide readers with an overview of population coverage, rather than to indicate the precise geographic origins of individual populations. The populations CDX, CHB, and CHS are displayed within the outline of China in alphabetical order, without any intention to indicate their exact geographic origin. We changed the respective figure caption to make this clear (page 19 of the revised manuscript):

“Map of the 888 samples from the 1000 Genomes project, mapping the samples to countries and not indicating detailed geographical origins of populations.”

Reviewer #3 (Public review):

This study successfully identified genetic loci associated with various traits by generating large-scale long-read sequencing data from a diverse set of samples. This study is significant because it not only produces large-scale long-read genome sequencing data but also demonstrates its application in actual genetics research. Given its potential utility in various fields, this study is expected to make a valuable contribution to the academic community and to this journal. However, there are several critical aspects that could be improved. Below are specific comments for consideration.

Strengths:

Producing high-quality, large-scale variant datasets and imputation datasets

Weaknesses:

(1) Data availability

Currently, it appears that only the Genomic Lens SV Panel is available on the webpage described in the Data Availability section. It is unclear whether the authors intend to release the raw sequencing data. Since the study utilized samples from the 1000 Genomes Project, there should be no restriction on making the data publicly accessible. Given this, would the authors consider making the raw sequencing reads publicly available? If so, NCBI SRA or EBI ENA would be the most appropriate repositories for data deposition. I strongly encourage the authors to consider public data release. Additionally, accessing the Genomic Lens SV Panel data does not seem straightforward. The manuscript should provide a more detailed description of how researchers can access and utilize these data. In my opinion, the best approach would be to upload the variant data (VCF files) to a public database such as the European Variation Archive (EVA) hosted by EBI.

I strongly request that the authors publicly deposit the variant data. At a minimum:

(a) The joint genotype data for all 888 samples from the 1000 Genomes Project must be publicly available.

Thank you for emphasising the importance of data sharing, which we agree with.

The Data Availability section of the manuscript already includes a link to https://opnme.com/genomiclens, where we make both the SV calls and the full and reduced SV imputation panels (provided as multi-sample VCF files) freely available. Based on the reviewer’s request, we now also reference the ENA repository project PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727), where the raw FASTQ files are available for download.

We have appended the Data Availability statement on page 22 of the revised manuscript as follows:

“Raw SV calls, the long-read sequencing-based SV imputation panel, and the SV summary statistics from 32 SV-wide association studies are available through the OpnMe initiative of Boehringer Ingelheim GmbH (https://opnme.com/genomiclens). The raw long-read sequencing data (FASTQ files) for the 1000 Genomes Project samples included in this study are accessible via the European Nucleotide Archive under accession number PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727). The dataset analysed here constitutes a subset of this broader collection.”

(b) For the UK Biobank samples, at least allele frequency data should be disclosed.

Supplementary Table 5 includes the allele frequencies of the SVs imputed into UK Biobank.

(c) Since eLife has a well-established data-sharing policy, compliance with these guidelines is essential for publication in this journal.

By sharing the FASTQ files, the SV calls, the SV imputation panels, the SV summary statistics, and (once processed by UK Biobank) the genotypes of SVs imputed into UK Biobank, we are providing all SV data generated in our study.

(2) Long-read sequencing data quality

While the manuscript presents N50 read length and mean or median read base quality for each sample in a table, it would be highly beneficial to visualize these data in figures as well. A violin plot or similar visualization summarizing these distributions would significantly improve data presentation.

Notably, the base quality of ONT long-read sequencing data appears lower than expected. This may be attributed to the use of pore version 9.4.1, but the unexpectedly low base quality still warrants attention. It would be helpful to include a small figure within Figure 2 to illustrate this point. A visual representation of read length distribution and base quality distribution would strengthen the manuscript.

We thank the reviewer for this suggestion. We have now included two violin plots (the new Supplementary Figure 1) to the revised manuscript, summarising a) the N50 read length per sequencing run and b) the median read quality per sequencing run. These plots provide a clearer visualisation of the underlying distributions. We do not consider the ONT base quality to be low. Importantly, structural variant detection is generally robust to modest variations of per-base quality. Therefore, we do not expect the observed base quality levels to significantly affect SV calling in this study.

(3) Variant detection precision, recall, and F1 score

This study focuses on insertions and deletions (indels) {greater than or equal to}50 bp, but it remains unclear how well variants <50 bp are detected. I am particularly interested in the precision, recall, and F1 score for variants between 5-49 bp.

While ONT base quality is relatively low, single-base variants are challenging to analyze, but variants {greater than or equal to}5 bp should still be detectable as their read accuracy is still approximately 90%, making analysis feasible. Given that Sniffles supports the detection of variants as small as 1 bp, I strongly encourage the authors to conduct an additional analysis.

A simple two-category classification (e.g., 5-49 bp and {greater than or equal to}50 bp) should suffice. Additionally, a comparative analysis with HiFi and short-read sequencing data would be highly valuable. If possible, I strongly recommend that all detected variants {greater than or equal to}5 bp be made publicly available as VCF files.

Because short InDels are available from high-coverage Illumina sequencing data generated for the same individuals (i.e., the data referred to as the NYGC dataset in our manuscript), we decided against calling such short variants from our lower coverage Oxford Nanopore data and thus concentrated our efforts on reliably calling longer variants covering at least 50 bp, consistent with the conventional definition of structural variants.

(4) Assembly-based methods

Given the low read accuracy and low sequencing depth in this dataset, it is understandable that genome assembly is challenging. However, the latest high-quality human genome datasets-such as those produced by the Human Pangenome Reference Consortium (HPRC)demonstrate that assembly-based approaches provide significant advantages, particularly for resolving complex and long structural variants.

Since HPRC data also utilize 1000 Genomes Project samples, it would be highly informative to compare the accuracy of ONT sequencing in this study with HPRC's assembly-based genome data. The recent publication on 47 HPRC samples provides a valuable reference for such a comparison. Given its relevance, the authors should consider providing a comparative analysis with HPRC data.

The aim of the present study was to generate an SV reference panel that enables SV imputation for large biobanks. Detailed assessments of ONT sequencing quality in general and comparisons to other sequencing efforts and technologies are out of scope for the present manuscript. We invite the scientific community to use the FASTQ files provided at ENA for conducting such detailed assessments in follow-up studies.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation