datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ms2prospect-ptms-ms2
PROSPECT PTMs - Fragment Ion Intensity Prediction (MS2)
A mass-spectrometry dataset for applied machine learning in proteomics, annotated, processed and split for the task of fragment ion intensity prediction.
Dataset Details
Curated by: Wilhelmlab - Technical University of Munich - School of Life Sciences - Germany
License: CC-BY4.0
Dataset Sources
The data is based on the PROSPECT PTMs datasets hosted in Zenodo [3][4][5][6].
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelmlab/prospect-ptms-ms2.ms2-peptide-replicate-retrieval
MS2 Peptide-Replicate-Retrieval Benchmark
A benchmark for evaluating spectrum-embedding models for tandem mass
spectrometry (MS2). It is a set of real experimental MS2 spectra, each labelled
with the peptide it was identified as (a peptide-spectrum match, PSM), pooled so
that every peptide is represented by many replicate acquisitions. The task:
does an embedding map replicate spectra of the same peptide close together?
15,649 spectra
1,000 unique peptide/charge labels (≈ 15… See the full description on the dataset page: https://huggingface.co/datasets/chrisagrams/ms2-peptide-replicate-retrieval.massive-v2-ms2-t095-l080-sharded-10gb
MassIVE v2 exact-MS2 training shards
This dataset is a training-oriented repack of
novogaia/massive-v2 at
revision 10c48d8184119829c48651b8a40ea5e0b9015687. It includes only source files ending in
_t0.95_l0.80_grouped.hdf5 and retains every row whose MS level is exactly 2.
Rows: 1,584,408,553
Eligible training rows: 1,516,329,213
Train shards: 65
Validation shards: 3
The training_eligible column records the canonical precursor, retention-time,
and usable-spectrum policy… See the full description on the dataset page: https://huggingface.co/datasets/novogaia/massive-v2-ms2-t095-l080-sharded-10gb.ms2_combinedms2_sparse_maxThis is a copy of the MS^2 dataset, except the input source documents of its validation split have been replaced by a sparse retriever. The retrieval pipeline used:
query: The background field of each example
corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract.
retriever: BM25 via PyTerrier with default settings
top-k strategy: "max", i.e. the number of documents retrieved, k, is set as the maximum number of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_sparse_max.tourism-datasetms2-multidoc-preprocessed-v3ms2_sparse_meanThis is a copy of the MS^2 dataset, except the input source documents of its validation split have been replaced by a sparse retriever. The retrieval pipeline used:
query: The background field of each example
corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract.
retriever: BM25 via PyTerrier with default settings
top-k strategy: "mean", i.e. the number of documents retrieved, k, is set as the mean number of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_sparse_mean.ms2_dense_maxThis is a copy of the MS^2 dataset, except the input source documents of its validation split have been replaced by a dense retriever. The retrieval pipeline used:
query: The background field of each example
corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract.
retriever: facebook/contriever-msmarco via PyTerrier with default settings
top-k strategy: "max", i.e. the number of documents retrieved, k, is set as… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_dense_max.MS2_trainms2_sparse_oracleThis is a copy of the MS^2 dataset, except the input source documents of its validation split have been replaced by a sparse retriever. The retrieval pipeline used:
query: The background field of each example
corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract.
retriever: BM25 via PyTerrier with default settings
top-k strategy: "oracle", i.e. the number of documents retrieved, k, is set as the original number… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_sparse_oracle.ms2-multidoc-preprocessed-v2Lms2_dense_meanThis is a copy of the MS^2 dataset, except the input source documents of its train, validation and test splits have been replaced by a dense retriever. The retrieval pipeline used:
query: The background field of each example
corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract.
retriever: facebook/contriever-msmarco via PyTerrier with default settings
top-k strategy: "max", i.e. the number of documents retrieved… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_dense_mean.ms2-multidoc-preprocessedms2_combined_valms2-multidoc-preprocessed-v2Hms2-preprocessed-multidocms2-cleaned-multidocMS2_testms2_dense_oracleThis is a copy of the MS^2 dataset, except the input source documents of the train, validation, and test splits have been replaced by a dense retriever.
query: The background field of each example
corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract.
retriever: facebook/contriever-msmarco via PyTerrier with default settings
top-k strategy: "oracle", i.e. the number of documents retrieved, k, is set as the… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_dense_oracle.timsTOF-ms2Prosit-2025-lac-ms2
Prosit lac 2025 - Fragment Ion Intensity Prediction (MS2)
A mass-spectrometry dataset for applied machine learning in proteomics, annotated, processed and split for the task of fragment ion intensity prediction.
Dataset Details
Curated by: Wilhelmlab - Technical University of Munich - School of Life Sciences - Germany
License: CC-BY4.0
Dataset Sources
The data is based on the PROSPECT dataset and on the Klac ChemIntelligence Library.
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelmlab/Prosit-2025-lac-ms2.Leader_Dataset_120725
Leader_Dataset_120725
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
tool-calling-conversations-ms2gn400
Tool Calling Conversations
An Arena-style dataset of anonymized, multi-turn conversations focused on real-world
tool use. It is intended for research, evaluation, and training of models that decide
when and how to call tools.
The conversations include:
Tool selection and no-tool decisions
Structured tool arguments
Sequential and parallel tool calls
Tool results and error recovery
Multi-step agent workflows
Final responses after tool execution
Data is organized into… See the full description on the dataset page: https://huggingface.co/datasets/dakr-pandas/tool-calling-conversations-ms2gn400.ms2_dataset_restructuredMS2_1shot_testms2525d-landunit-symbols-csvautoeval-staging-eval-project-ben-yu__ms2_combined-823f066f-12515671
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: Blaise-g/long_t5_global_large_pubmed_explanatory
Dataset: ben-yu/ms2_combined
Config: ben-yu--ms2_combined
Split: train
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @ben-yu for evaluating this model.
CSU-MS2-DB
