CoolFace
Datasetpublic

TopRepo/C-elegans

TopRepo is a top-down spectral repository containing more than 12 million MS/MS spectra from 12 species. For each MS raw file in TopRepo, msconvert is used to convert the raw file to a centroided mzML file, and then TopFD is employed to deconvolute spectra in the centroided mzML file to one or several msalign files and proteoform feature files. The msalign files are searched against its corresponding proteome sequence database for spectral identification using TopPIC. The identification… See the full description on the dataset page: https://huggingface.co/datasets/TopRepo/C-elegans.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes20downloads
Dataset Card

TopRepo is a top-down spectral repository containing more than 12 million MS/MS spectra from 12 species. For each MS raw file in TopRepo, msconvert is used to convert the raw file to a centroided mzML file, and then TopFD is employed to deconvolute spectra in the centroided mzML file to one or several msalign files and proteoform feature files. The msalign files are searched against its corresponding proteome sequence database for spectral identification using TopPIC. The identification results are stored in TSV files.

Using mzML files, msalign files, feature files, and spectral identification (TSV) files generated from the data analysis pipeline, Python scripts in this repository are used to generate TSV files with comprehensive spectral information, annotated msalign files, and annotated mgf files.

Dataset Structure

The dataset consists of two TSV files for the species Caenorhabditis elegans.

  1. 1.Metadata Table toprepo_c_elegans_meta_table_v1.2.0.tsv Contains experiment-level metadata including:
  2. 2.dataset identifiers
  3. 3.species information
  4. 4.instrument metadata
  5. 5.dissociation methods
  6. 6.counts of spectra, proteins, and proteoforms
  1. 1.Spectrum Table toprepo_c_elegans_spectrum_table_ms2_v1.2.0.tsv Contains spectrum-level MS/MS analytical data including:
  2. 2.scan identifiers
  3. 3.precursor ion information
  4. 4.fragmentation measurements
  5. 5.proteoform annotations
  6. 6.protein identifications
  7. 7.statistical confidence metrics
  8. 8.Annotated .msalign file
  9. 9.Annotated .mgf file

Sources

  • Repository https://toprepo.org/
  • Paper https://www.biorxiv.org/content/10.64898/2026.02.20.707032v1
  • Github https://github.com/toppic-suite/toprepo

Contact

For questions regarding this dataset, please contact the TopRepo team through: xwliu@tulane.edu kli7@tulane.edu

Copyright (c) 2025 - 2026, Tulane University.