MARATTO

dataset · Zenodo (CERN European Organization for Nuclear Research)

Curated carbapenemase protein sequence dataset and reproducibility files

In plain language

A curated protein sequence dataset and associated reproducibility files have been assembled to support the integrated analysis of carbapenemases. The initial collection of 4,178 protein accession records was dereplicated into 3,349 exact unique amino acid sequences. Multi-method validation using CARD RGI, NCBI AMRFinderPlus, Pfam domains, catalytic motifs, and phylogenetics confirmed 3,246 unique proteins across 4,069 accessions. From this, a primary quality-controlled cohort of 2,548 full-length unique proteins was established alongside a secondary cohort of 698 unique proteins. The deposited resource provides sequence mappings, cluster information, taxonomic and temporal records, phylogenetic trees, and analysis scripts, with all source entries remaining fully traceable to UniProtKB and UniParc identifiers.

Key takeaways

  • A source collection of 4,178 carbapenemase accession records was dereplicated into 3,349 unique amino acid sequences.
  • Integrated validation using established annotation tools, catalytic motifs, and phylogenetics retained 3,246 unique proteins.
  • The primary quality-controlled cohort contains 2,548 full-length unique proteins representing 3,294 accession records.
  • The open archive supplies complete analytical scripts, alignments, phylogenetic trees, and database mappings traceable to UniProtKB and UniParc.
  • Reported taxonomic and temporal metrics reflect database curation records rather than real-world biological prevalence or isolate collection dates.

Why it matters

Carbapenemase enzymes cause resistance to critical antibiotics, making accurate identification essential. Public databases often contain duplicate entries or inconsistent annotations. By systematically validating sequences and establishing a high-confidence, non-redundant reference cohort, this resource helps researchers rely on standardised, reproducible genomic and protein data when studying antimicrobial resistance mechanisms.

Commercialisation angle

The abstract does not indicate a commercial application pathway, presenting instead a curated reference dataset and reproducibility files for academic and bioinformatics research.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

This dataset contains curated protein-sequence data and reproducibility materials supporting an integrated analysis of carbapenemase annotation confidence, exact-sequence redundancy, phylogenetic family/lineage resolution, taxonomic representation, and temporal database representation. The source collection comprised 4,178 protein accession records, dereplicated into 3,349 exact unique amino-acid sequences. Integrated validation using CARD RGI, NCBI AMRFinderPlus, Pfam-domain evidence, class-specific catalytic motifs, and phylogenetic analyses retained 3,246 unique proteins representing 4,069 accession records. The primary analytical cohort comprised 2,548 full-length QC-passing unique proteins representing 3,294 accession records, while the supported secondary cohort comprised 698 unique proteins representing 775 accession records. The deposited archive includes accession-to-sequence mappings, exact-sequence clusters, cohort assignments, integrated annotation and validation outputs, final family/lineage assignments, phylogenetic alignments and trees, taxonomic and temporal analysis tables, Supplementary Tables S1–S20, analysis scripts, and reproducibility records. Source protein records remain traceable to their public UniProtKB and UniParc accessions. Taxonomic counts describe representation within the curated dataset and should not be interpreted as prevalence, transmission frequency, or host specificity. Temporal variables correspond to database-record creation/deposition dates and should not be interpreted as isolate-collection dates, biological emergence, or discovery dates.

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.5281/zenodo.21882475

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.