CoolFace
Datasetpublic

Wilhelmlab/Prosit-2025-lac-ms2

Prosit lac 2025 - Fragment Ion Intensity Prediction (MS2) A mass-spectrometry dataset for applied machine learning in proteomics, annotated, processed and split for the task of fragment ion intensity prediction. Dataset Details Curated by: Wilhelmlab - Technical University of Munich - School of Life Sciences - Germany License: CC-BY4.0 Dataset Sources The data is based on the PROSPECT dataset and on the Klac ChemIntelligence Library. Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelmlab/Prosit-2025-lac-ms2.

sourceHugging Facecc-by-4.0updated 11mo agoView on Hugging Face
0likes17downloads
Dataset Card

Prosit lac 2025 - Fragment Ion Intensity Prediction (MS2)

A mass-spectrometry dataset for applied machine learning in proteomics, annotated, processed and split for the task of fragment ion intensity prediction.

Dataset Details

  • —Curated by: Wilhelmlab - Technical University of Munich - School of Life Sciences - Germany
  • —License: CC-BY4.0

Dataset Sources

The data is based on the PROSPECT dataset and on the Klac ChemIntelligence Library.

  • —Repository: https://github.com/wilhelm-lab/PROSPECT

Uses

The dataset is intended to be used for fragment ion intensity prediction given a peptide sequence.

Use the following lines to load the data (train/val/test):

python
# main data for training and evaluation; contains train, validation, test splits
main_dataset = load_dataset("Wilhelmlab/Prosit-2025-lac-ms2")

Dataset Creation

Curation Rationale

The dataset is intended to serve as a reference benchmark dataset for fragment ion intensity prediction, processed, split, and ready-to-use for developing deep learning models for this specific task.

Source Data

The upstream source data is based on the ProteomeTools datasets available on PRIDE [1][2].

Data Collection and Processing

[More Information Needed]

Annotations

The annotations are based on an expert system [7] with a set of rules listed in the PROSPECT paper [8]. The vector of intensities is collected in one column named intensities_raw.

Personal and Sensitive Information

The dataset does not contain any personal, sensitive, or private data.

Recommendations

We recommend using the holdout configuration for solely evaluation models at the end of the research iteration.

Citation

BibTeX:

[More Information Needed]

APA:

References

[1] Daniel P Zolg, Mathias Wilhelm, Karsten Schnatbaum, Johannes Zerweck, Tobias Knaute, Bernard Delanghe, Derek J Bailey, Siegfried Gessulat, Hans-Christian Ehrlich, Maximilian Weininger, et al. Building proteometools based on a complete synthetic human proteome. Nature methods, 14(3):259–262, 2017.

[2] Nadin Neuhauser, Annette Michalski, Jürgen Cox, and Matthias Mann. Expert system for computer-assisted annotation of ms/ms spectra. Molecular & Cellular Proteomics, 11(11):1500– 1509, 2012.

[3] Omar Shouman, Wassim Gabriel, Victor-George Giurcoiu, Vitor Sternlicht, and Mathias Wil- helm. PROSPECT: Labeled tandem mass spectrometry dataset for machine learning in pro- teomics. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 32882–32896. Curran Associates, Inc., 2022.

Dataset Card Contact

mathias.wilhelm@tum.de

Wilhelmlab, TU Munich, School of Life Sciences, Germany.