datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LRGB_Peptides-func
LRGB Peptides-func
Peptides-func (Peptides functional) dataset, part of Long Range Graph Benchmark (LRGB) [1]. It is intended to be used through
scikit-fingerprints library.
The task is to predict functional properties of peptides.
Characteristic
Description
Tasks
10
Task type
classification
Total samples
15535
Recommended split
stratified random
Recommended metric
AUPRC
References
[1]
Dwivedi, Vijay Prakash, et al.
"Long Range Graph Benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/LRGB_Peptides-func.LRGB_Peptides-struct
LRGB Peptides-struct
Peptides-struct (Peptides structural) dataset, part of Long Range Graph Benchmark (LRGB) [1]. It is intended to be used through
scikit-fingerprints library.
The task is to predict structural properties of peptides. Note that this is raw data, whereas the original paper [1] specifies
that targets should be standardized (mean 0, standard deviation 1) before training and evaluation. scikit-fingerprints does this by
default in the loader function, otherwise this… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/LRGB_Peptides-struct.ms2-peptide-replicate-retrieval
MS2 Peptide-Replicate-Retrieval Benchmark
A benchmark for evaluating spectrum-embedding models for tandem mass
spectrometry (MS2). It is a set of real experimental MS2 spectra, each labelled
with the peptide it was identified as (a peptide-spectrum match, PSM), pooled so
that every peptide is represented by many replicate acquisitions. The task:
does an embedding map replicate spectra of the same peptide close together?
15,649 spectra
1,000 unique peptide/charge labels (≈ 15… See the full description on the dataset page: https://huggingface.co/datasets/chrisagrams/ms2-peptide-replicate-retrieval.peptideclm-2-pretraining-data
PeptideMTR Training Data
This repository contains the dataset for the PeptideMTR paper. It is designed for SMILES encoder models trained by masked-language modeling (MLM) and/or multi-target regression (MTR) tasks, focusing on mapping peptide sequences to biochemical properties.
Link to the manuscript will be added here when available.
Dataset Summary
The dataset includes peptide sequences paired with 99 RDKit-derived descriptors representing various physicochemical… See the full description on the dataset page: https://huggingface.co/datasets/aaronfeller/peptideclm-2-pretraining-data.peptide-reasoning-benchmark
Peptide Reasoning Benchmark
PEB v1.0-RC benchmark release for peptide-reasoning model evaluation.
Includes cases, splits, baselines, references, and leaderboard artifacts.
GitHub: https://github.com/ray-r-ren/peptide-reasoning-bench
Trained a small reference LoRA model: https://huggingface.co/rayrren/the-spice-v0-mvp
bbb-peptides
BBB Peptide Dataset (TFG)
Curated blood–brain barrier (BBB) permeability dataset for peptide sequences with Boltz-predicted 3D structures and physicochemical descriptors.
Summary
Field
Value
Rows
825
With structure
825
BBB+
410
BBB−
415
Variant
full
Files
peptides.parquet — one row per peptide: sequence, label, splits, physicochemical features, structure quality metrics, relative structure paths… See the full description on the dataset page: https://huggingface.co/datasets/manumartinm/bbb-peptides.peptideforge-dataset
PeptideForge Dataset
Dataset Description
This dataset repo packages the processed training, validation, and test splits used by the
PeptideForge project for conditioned peptide generation and AMP scoring.
It exposes three Hub configs:
config
purpose
splits
generator_text
Conditioned text corpus exported as parsed CSV rows
train / validation / test
generator_structured
Structured generator table with features and conditioned prompts
train / validation / test… See the full description on the dataset page: https://huggingface.co/datasets/HakimT/peptideforge-dataset.glp1-peptide-dosage-benchmarks-2026
Clinical Peptide Reconstitution & Micro-Dosing Math
Empirical benchmark dataset by Groundwork Research (https://gworky.com).
Full interactive decision engine available at: https://gworky.com/tools/peptide-reconstitution-calculator.
Description
Precision microgram-to-unit reconstitution calculator for GLP-1 analogues, Tirzepatide, and research peptides with syringe dead-space calibration.
Primary source authority: https://gworky.com/body
fda-peptide-human-evidence
Seven FDA-reviewed peptides: claims coded for identity, administration, outcome and replication
A claim-level comparison of the seven peptide pairs reviewed at the FDA Pharmacy Compounding Advisory Committee meeting of 23-24 July 2026. The data separates molecular identity, administration to people, claimed outcomes and independent replication.
Read the evidence-led article: https://lifesco.re/edge/which-peptide-claims-have-actually-been-tested-in-people/
Archived version and… See the full description on the dataset page: https://huggingface.co/datasets/lifescore/fda-peptide-human-evidence.peptide-compound-reference
Peptide Compound Reference Database
Published by Dosi Health
Open dataset of 185 research peptides, GLP-1 agonists, and growth-hormone secretagogues — with half-life, default dosing, injection routes, and compound metadata.
This dataset powers the compound library inside Dosi, a free peptide / GLP-1 / TRT tracker available on iOS, Android, and web.
Files
peptides.json — full dataset, 185 entries, all fields
peptides.csv — flat CSV of the most-used columns… See the full description on the dataset page: https://huggingface.co/datasets/hdeanwhitebot/peptide-compound-reference.research-peptides-reference
Research Peptides Reference Dataset
A clean, machine-readable reference table of research-grade peptides commonly
discussed in biochemistry and drug-discovery literature. Each entry combines a
curated research category with verified physicochemical properties pulled from
PubChem (PUG REST): PubChem CID, molecular
formula, molecular weight, canonical SMILES and IUPAC name.
The goal is a small, high-signal starting point for cheminformatics, tabular ML,
educational tooling and… See the full description on the dataset page: https://huggingface.co/datasets/PeptidosSuplementos/research-peptides-reference.peptide_UniRef50_0_50_with_smiles_clean_tokenized
