datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prospect-ptms-ms2
PROSPECT PTMs - Fragment Ion Intensity Prediction (MS2)
A mass-spectrometry dataset for applied machine learning in proteomics, annotated, processed and split for the task of fragment ion intensity prediction.
Dataset Details
Curated by: Wilhelmlab - Technical University of Munich - School of Life Sciences - Germany
License: CC-BY4.0
Dataset Sources
The data is based on the PROSPECT PTMs datasets hosted in Zenodo [3][4][5][6].
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelmlab/prospect-ptms-ms2.ms2-peptide-replicate-retrieval
MS2 Peptide-Replicate-Retrieval Benchmark
A benchmark for evaluating spectrum-embedding models for tandem mass
spectrometry (MS2). It is a set of real experimental MS2 spectra, each labelled
with the peptide it was identified as (a peptide-spectrum match, PSM), pooled so
that every peptide is represented by many replicate acquisitions. The task:
does an embedding map replicate spectra of the same peptide close together?
15,649 spectra
1,000 unique peptide/charge labels (≈ 15… See the full description on the dataset page: https://huggingface.co/datasets/chrisagrams/ms2-peptide-replicate-retrieval.ms2_sparse_maxThis is a copy of the MS^2 dataset, except the input source documents of its validation split have been replaced by a sparse retriever. The retrieval pipeline used:
query: The background field of each example
corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract.
retriever: BM25 via PyTerrier with default settings
top-k strategy: "max", i.e. the number of documents retrieved, k, is set as the maximum number of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_sparse_max.ms2_dense_maxThis is a copy of the MS^2 dataset, except the input source documents of its validation split have been replaced by a dense retriever. The retrieval pipeline used:
query: The background field of each example
corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract.
retriever: facebook/contriever-msmarco via PyTerrier with default settings
top-k strategy: "max", i.e. the number of documents retrieved, k, is set as… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_dense_max.ms2-multidoc-preprocessed-v3ms2_sparse_meanThis is a copy of the MS^2 dataset, except the input source documents of its validation split have been replaced by a sparse retriever. The retrieval pipeline used:
query: The background field of each example
corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract.
retriever: BM25 via PyTerrier with default settings
top-k strategy: "mean", i.e. the number of documents retrieved, k, is set as the mean number of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_sparse_mean.MS2_trainms2-multidoc-preprocessed-v2Lms2_sparse_oracleThis is a copy of the MS^2 dataset, except the input source documents of its validation split have been replaced by a sparse retriever. The retrieval pipeline used:
query: The background field of each example
corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract.
retriever: BM25 via PyTerrier with default settings
top-k strategy: "oracle", i.e. the number of documents retrieved, k, is set as the original number… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_sparse_oracle.ms2-multidoc-preprocessedms2_dense_meanThis is a copy of the MS^2 dataset, except the input source documents of its train, validation and test splits have been replaced by a dense retriever. The retrieval pipeline used:
query: The background field of each example
corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract.
retriever: facebook/contriever-msmarco via PyTerrier with default settings
top-k strategy: "max", i.e. the number of documents retrieved… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_dense_mean.ms2-multidoc-preprocessed-v2Hms2-cleaned-multidocms2-preprocessed-multidocms2_dense_oracleThis is a copy of the MS^2 dataset, except the input source documents of the train, validation, and test splits have been replaced by a dense retriever.
query: The background field of each example
corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract.
retriever: facebook/contriever-msmarco via PyTerrier with default settings
top-k strategy: "oracle", i.e. the number of documents retrieved, k, is set as the… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_dense_oracle.MS2_testtimsTOF-ms2Prosit-2025-lac-ms2
Prosit lac 2025 - Fragment Ion Intensity Prediction (MS2)
A mass-spectrometry dataset for applied machine learning in proteomics, annotated, processed and split for the task of fragment ion intensity prediction.
Dataset Details
Curated by: Wilhelmlab - Technical University of Munich - School of Life Sciences - Germany
License: CC-BY4.0
Dataset Sources
The data is based on the PROSPECT dataset and on the Klac ChemIntelligence Library.
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelmlab/Prosit-2025-lac-ms2.MS2_1shot_testtool-calling-conversations-ms2gn400
Tool Calling Conversations
An Arena-style dataset of anonymized, multi-turn conversations focused on real-world
tool use. It is intended for research, evaluation, and training of models that decide
when and how to call tools.
The conversations include:
Tool selection and no-tool decisions
Structured tool arguments
Sequential and parallel tool calls
Tool results and error recovery
Multi-step agent workflows
Final responses after tool execution
Data is organized into… See the full description on the dataset page: https://huggingface.co/datasets/dakr-pandas/tool-calling-conversations-ms2gn400.ms2525d-landunit-symbols-csvautoeval-staging-eval-project-ben-yu__ms2_combined-823f066f-12515671
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: Blaise-g/long_t5_global_large_pubmed_explanatory
Dataset: ben-yu/ms2_combined
Config: ben-yu--ms2_combined
Split: train
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @ben-yu for evaluating this model.
ms2525d-landunit-symbols-csv
