datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LRGB_Peptides-func
LRGB Peptides-func
Peptides-func (Peptides functional) dataset, part of Long Range Graph Benchmark (LRGB) [1]. It is intended to be used through
scikit-fingerprints library.
The task is to predict functional properties of peptides.
Characteristic
Description
Tasks
10
Task type
classification
Total samples
15535
Recommended split
stratified random
Recommended metric
AUPRC
References
[1]
Dwivedi, Vijay Prakash, et al.
"Long Range Graph Benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/LRGB_Peptides-func.LRGB_Peptides-struct
LRGB Peptides-struct
Peptides-struct (Peptides structural) dataset, part of Long Range Graph Benchmark (LRGB) [1]. It is intended to be used through
scikit-fingerprints library.
The task is to predict structural properties of peptides. Note that this is raw data, whereas the original paper [1] specifies
that targets should be standardized (mean 0, standard deviation 1) before training and evaluation. scikit-fingerprints does this by
default in the loader function, otherwise this… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/LRGB_Peptides-struct.peptideforge-dataset
PeptideForge Dataset
Dataset Description
This dataset repo packages the processed training, validation, and test splits used by the
PeptideForge project for conditioned peptide generation and AMP scoring.
It exposes three Hub configs:
config
purpose
splits
generator_text
Conditioned text corpus exported as parsed CSV rows
train / validation / test
generator_structured
Structured generator table with features and conditioned prompts
train / validation / test… See the full description on the dataset page: https://huggingface.co/datasets/HakimT/peptideforge-dataset.fda-peptide-human-evidence
Seven FDA-reviewed peptides: claims coded for identity, administration, outcome and replication
A claim-level comparison of the seven peptide pairs reviewed at the FDA Pharmacy Compounding Advisory Committee meeting of 23-24 July 2026. The data separates molecular identity, administration to people, claimed outcomes and independent replication.
Read the evidence-led article: https://lifesco.re/edge/which-peptide-claims-have-actually-been-tested-in-people/
Archived version and… See the full description on the dataset page: https://huggingface.co/datasets/lifescore/fda-peptide-human-evidence.research-peptides-reference
Research Peptides Reference Dataset
A clean, machine-readable reference table of research-grade peptides commonly
discussed in biochemistry and drug-discovery literature. Each entry combines a
curated research category with verified physicochemical properties pulled from
PubChem (PUG REST): PubChem CID, molecular
formula, molecular weight, canonical SMILES and IUPAC name.
The goal is a small, high-signal starting point for cheminformatics, tabular ML,
educational tooling and… See the full description on the dataset page: https://huggingface.co/datasets/PeptidosSuplementos/research-peptides-reference.
