ms2
Datasets
All datasets matching “ms2”ms2prospect-ptms-ms2
PROSPECT PTMs - Fragment Ion Intensity Prediction (MS2)
A mass-spectrometry dataset for applied machine learning in proteomics, annotated, processed and split for the task of fragment ion intensity prediction.
Dataset Details
Curated by: Wilhelmlab - Technical University of Munich - School of Life Sciences - Germany
License: CC-BY4.0
Dataset Sources
The data is based on the PROSPECT PTMs datasets hosted in Zenodo [3][4][5][6].
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelmlab/prospect-ptms-ms2.ms2-peptide-replicate-retrieval
MS2 Peptide-Replicate-Retrieval Benchmark
A benchmark for evaluating spectrum-embedding models for tandem mass
spectrometry (MS2). It is a set of real experimental MS2 spectra, each labelled
with the peptide it was identified as (a peptide-spectrum match, PSM), pooled so
that every peptide is represented by many replicate acquisitions. The task:
does an embedding map replicate spectra of the same peptide close together?
15,649 spectra
1,000 unique peptide/charge labels (≈ 15… See the full description on the dataset page: https://huggingface.co/datasets/chrisagrams/ms2-peptide-replicate-retrieval.massive-v2-ms2-t095-l080-sharded-10gb
MassIVE v2 exact-MS2 training shards
This dataset is a training-oriented repack of
novogaia/massive-v2 at
revision 10c48d8184119829c48651b8a40ea5e0b9015687. It includes only source files ending in
_t0.95_l0.80_grouped.hdf5 and retains every row whose MS level is exactly 2.
Rows: 1,584,408,553
Eligible training rows: 1,516,329,213
Train shards: 65
Validation shards: 3
The training_eligible column records the canonical precursor, retention-time,
and usable-spectrum policy… See the full description on the dataset page: https://huggingface.co/datasets/novogaia/massive-v2-ms2-t095-l080-sharded-10gb.ms2_combinedms2_sparse_maxThis is a copy of the MS^2 dataset, except the input source documents of its validation split have been replaced by a sparse retriever. The retrieval pipeline used:
query: The background field of each example
corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract.
retriever: BM25 via PyTerrier with default settings
top-k strategy: "max", i.e. the number of documents retrieved, k, is set as the maximum number of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_sparse_max.
