Wilhelmlab/Prosit-2025-lac-ms2
Prosit lac 2025 - Fragment Ion Intensity Prediction (MS2) A mass-spectrometry dataset for applied machine learning in proteomics, annotated, processed and split for the task of fragment ion intensity prediction. Dataset Details Curated by: Wilhelmlab - Technical University of Munich - School of Life Sciences - Germany License: CC-BY4.0 Dataset Sources The data is based on the PROSPECT dataset and on the Klac ChemIntelligence Library. Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelmlab/Prosit-2025-lac-ms2.
Prosit lac 2025 - Fragment Ion Intensity Prediction (MS2)
A mass-spectrometry dataset for applied machine learning in proteomics, annotated, processed and split for the task of fragment ion intensity prediction.
Dataset Details
- Curated by: Wilhelmlab - Technical University of Munich - School of Life Sciences - Germany
- License: CC-BY4.0
Dataset Sources
The data is based on the PROSPECT dataset and on the Klac ChemIntelligence Library.
- Repository: https://github.com/wilhelm-lab/PROSPECT
Uses
The dataset is intended to be used for fragment ion intensity prediction given a peptide sequence.
Use the following lines to load the data (train/val/test):
# main data for training and evaluation; contains train, validation, test splits
main_dataset = load_dataset("Wilhelmlab/Prosit-2025-lac-ms2")Dataset Creation
Curation Rationale
The dataset is intended to serve as a reference benchmark dataset for fragment ion intensity prediction, processed, split, and ready-to-use for developing deep learning models for this specific task.
Source Data
The upstream source data is based on the ProteomeTools datasets available on PRIDE [1][2].
Data Collection and Processing
[More Information Needed]
Annotations
The annotations are based on an expert system [7] with a set of rules listed in the PROSPECT paper [8]. The vector of intensities is collected in one column named intensities_raw.
Personal and Sensitive Information
The dataset does not contain any personal, sensitive, or private data.
Recommendations
We recommend using the holdout configuration for solely evaluation models at the end of the research iteration.
Citation
BibTeX:
[More Information Needed]
APA:
References
[1] Daniel P Zolg, Mathias Wilhelm, Karsten Schnatbaum, Johannes Zerweck, Tobias Knaute, Bernard Delanghe, Derek J Bailey, Siegfried Gessulat, Hans-Christian Ehrlich, Maximilian Weininger, et al. Building proteometools based on a complete synthetic human proteome. Nature methods, 14(3):259–262, 2017.
[2] Nadin Neuhauser, Annette Michalski, Jürgen Cox, and Matthias Mann. Expert system for computer-assisted annotation of ms/ms spectra. Molecular & Cellular Proteomics, 11(11):1500– 1509, 2012.
[3] Omar Shouman, Wassim Gabriel, Victor-George Giurcoiu, Vitor Sternlicht, and Mathias Wil- helm. PROSPECT: Labeled tandem mass spectrometry dataset for machine learning in pro- teomics. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 32882–32896. Curran Associates, Inc., 2022.
Dataset Card Contact
mathias.wilhelm@tum.de
Wilhelmlab, TU Munich, School of Life Sciences, Germany.
