MARATTO

article · Genome Research

High-coverage nanopore sequencing of samples from the 1000 Genomes Project to build a comprehensive catalog of human genetic variation

202462 citationsOpen accessUniversity of the Witwatersrand

In plain language

Fewer than half of individuals suspected of having a monogenic condition receive a precise molecular diagnosis through routine clinical testing. While long-read sequencing could improve testing, clinical adoption has been limited by the lack of baseline control datasets needed to filter and prioritise variants. To address this limitation, high-coverage nanopore sequencing was applied to an initial set of 100 diverse human genomes representing five global superpopulations. The approach identified an average of 24,543 structural variants per genome, revealing gene-disrupting variants and pathogenic repeat expansions that were missed by short-read sequencing technologies. The sequencing data also captured methylation profiles, identifying novel differentially methylated regions alongside expected patterns at imprinted loci and skewed X-inactivation. All raw and processed data are publicly accessible to provide a reference resource for discovering pathogenic structural variants.

Key takeaways

  • Nanopore long-read sequencing of 100 diverse human samples revealed an average of 24,543 high-confidence structural variants per genome.
  • The analysis detected pathogenic repeat expansions and gene-disrupting structural variants that were not identified by standard short-read methods.
  • Sequencing data captured methylation profiles across the genome, detecting novel differentially methylated regions as well as expected patterns at imprinted loci.
  • All raw and processed data have been released publicly to serve as a baseline control resource for clinical variant filtering and prioritisation.

Why it matters

Most individuals tested for suspected monogenic disorders fail to receive a precise molecular diagnosis. Long-read sequencing captures complex genomic alterations that standard methods miss, but clinical interpretation requires reference datasets from diverse populations to separate harmless variation from disease-causing mutations. Making these high-coverage datasets publicly available gives diagnostic teams the control data required to interpret complex structural variants accurately.

Commercialisation angle

The release of this public reference dataset directly supports clinical diagnostics laboratories and genomic software developers. It enables the tertiary analysis and filtering of long-read sequencing data used to diagnose suspected monogenic conditions. As the dataset and summary statistics are already publicly available, the resource is immediately ready for integration into clinical bioinformatics pipelines and variant prioritisation tools without further commercial translation steps.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Fewer than half of individuals with a suspected Mendelian or monogenic condition receive a precise molecular diagnosis after comprehensive clinical genetic testing. Improvements in data quality and costs have heightened interest in using long-read sequencing (LRS) to streamline clinical genomic testing, but the absence of control data sets for variant filtering and prioritization has made tertiary analysis of LRS data challenging. To address this, the 1000 Genomes Project (1KGP) Oxford Nanopore Technologies Sequencing Consortium aims to generate LRS data from at least 800 of the 1KGP samples. Our goal is to use LRS to identify a broader spectrum of variation so we may improve our understanding of normal patterns of human variation. Here, we present data from analysis of the first 100 samples, representing all 5 superpopulations and 19 subpopulations. These samples, sequenced to an average depth of coverage of 37× and sequence read N50 of 54 kbp, have high concordance with previous studies for identifying single nucleotide and indel variants outside of homopolymer regions. Using multiple structural variant (SV) callers, we identify an average of 24,543 high-confidence SVs per genome, including shared and private SVs likely to disrupt gene function as well as pathogenic expansions within disease-associated repeats that were not detected using short reads. Evaluation of methylation signatures revealed expected patterns at known imprinted loci, samples with skewed X-inactivation patterns, and novel differentially methylated regions. All raw sequencing data, processed data, and summary statistics are publicly available, providing a valuable resource for the clinical genetics community to discover pathogenic SVs.

Research topics

  • Genetic Syndromes and Imprinting
  • Genomics and Rare Diseases
  • Cancer Genomics and Diagnostics

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1101/gr.279273.124

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.