datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MoleculeNet_BACE
MoleculeNet BACE
BACE dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict binding results for a set of inhibitors of humanβ-secretase 1 (BACE-1).
Characteristic
Description
Tasks
1
Task type
classification
Total samples
1513
Recommended split
scaffold
Recommended metricAUROC
References
[1]
Govindan Subramanian et al.
"Computational Modeling of β-Secretase 1 (BACE-1)… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_BACE.MoleculeNet_PCBA
MoleculeNet PCBA
PCBA (PubChem BioAssay) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict biological activity against 128 bioassays, generated by high-throughput screening (HTS). All tasks are binary active/non-active.
Note that targets have missing values. Algorithms should be evaluated only on present labels. For training data, you may want to impute them, e.g. with zeros.
Characteristic… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_PCBA.LRGB_Peptides-func
LRGB Peptides-func
Peptides-func (Peptides functional) dataset, part of Long Range Graph Benchmark (LRGB) [1]. It is intended to be used through
scikit-fingerprints library.
The task is to predict functional properties of peptides.
Characteristic
Description
Tasks
10
Task type
classification
Total samples
15535
Recommended split
stratified random
Recommended metric
AUPRC
References
[1]
Dwivedi, Vijay Prakash, et al.
"Long Range Graph Benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/LRGB_Peptides-func.MoleculeNet_Tox21
MoleculeNet Tox21
Tox21 dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict 12 toxicity targets, including nuclear receptors and stress response pathways. All tasks are binary.
Note that targets have missing values. Algorithms should be evaluated only on present labels. For training data, you may want to impute them, e.g. with zeros.
Characteristic
Description
Tasks
12
Task type
multitask… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_Tox21.MoleculeNet_ESOL
MoleculeNet ESOL
ESOL (Estimated SOLubility) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict aqueous solubility. Targets are log-transformed, and the unit is log mols per litre (log Mol/L).
Characteristic
Description
Tasks
1
Task type
regression
Total samples
1128
Recommended split
scaffold
Recommended metric
RMSE
References
[1]
John S. Delaney
"ESOL: Estimating… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ESOL.MoleculeNet_BBBP
MoleculeNet BBBP
BBBP (Blood-Brain Barrier Penetration) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict blood-brain barrier penetration (barrier permeability) of small drug-like molecules.
Characteristic
Description
Tasks
1
Task type
classification
Total samples
2039
Recommended split
scaffold
Recommended metric
AUROC
References
[1]
Ines Filipa Martins et al.
"A… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_BBBP.MoleculeNet_ClinTox
MoleculeNet ClinTox
Load and return the ClinTox dataset, part of MoleculeNet [1] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict drug approval viability, by predicting clinical trial toxicity and final FDA approval status. Both tasks are binary.
Characteristic
Description
Tasks
2
Task type
multitask classification
Total samples
1477
Recommended split
scaffold
Recommended metric
AUROC
References
[1]… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ClinTox.MoleculeNet_SIDER
MoleculeNet SIDER
Load and return the SIDER (Side Effect Resource) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict adverse drug reactions (ADRs) as drug side effects to 27 system organ classes in MedDRA classification. All tasks are binary.
Characteristic
Description
Tasks
12
Task type
multitask classification
Total samples
7831
Recommended split
scaffold
Recommended metric… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_SIDER.MoleculeNet_Lipophilicity
MoleculeNet Lipophilicity
Lipophilicity dataset, part of MoleculeNet [1] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict octanol/water distribution coefficient (logD) at pH 7.4. Targets are already log transformed, and are a unitless ratio.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
4200
Recommended split
scaffold
Recommended metric
RMSE
References
[1]
Wu, Zhenqin, et al.… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_Lipophilicity.MoleculeNet_MUV
MoleculeNet MUV
Load and return the MUV (Maximum Unbiased Validation) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict 17 targets designed for validation of virtual screening techniques, based on PubChem BioAssays. All tasks are binary.
Note that targets have missing values. Algorithms should be evaluated only on present labels. For training data, you may want to impute them, e.g. with zeros.… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_MUV.MoleculeNet_ToxCast
MoleculeNet ToxCast
ToxCast dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict 617 toxicity targets from a large library of compounds based on in vitro high-throughput screening. All tasks are binary.
Note that targets have missing values. Algorithms should be evaluated only on present labels. For training data, you may want to impute them, e.g. with zeros.
Characteristic
Description
Tasks… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ToxCast.MoleculeNet_HIV
MoleculeNet HIV
HIV dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict ability of molecules to inhibit HIV replication.
Characteristic
Description
Tasks
1
Task type
classification
Total samples
41127
Recommended split
scaffold
Recommended metric
AUROCWarning: in newer RDKit vesions, 7 molecules from the original dataset are not read correctly due to disallowed
hypervalent states of… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_HIV.MoleculeNet_FreeSolv
MoleculeNet FreeSolv
FreeSolv (Free Solvation Database) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict hydration free energy of small molecules in water. Targets are in kcal/mol.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
642
Recommended split
scaffold
Recommended metric
RMSE
References
[1]
Mobley, D.L., Guthrie, J.P.
"FreeSolv: a… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_FreeSolv.LRGB_Peptides-struct
LRGB Peptides-struct
Peptides-struct (Peptides structural) dataset, part of Long Range Graph Benchmark (LRGB) [1]. It is intended to be used through
scikit-fingerprints library.
The task is to predict structural properties of peptides. Note that this is raw data, whereas the original paper [1] specifies
that targets should be standardized (mean 0, standard deviation 1) before training and evaluation. scikit-fingerprints does this by
default in the loader function, otherwise this… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/LRGB_Peptides-struct.LLM-fingerprinted-adaptervideo-fingerprinting-tvsumlitmusLLM-fingerprinted-SFTLLM-fingerprintedExpansionRx_OpenADMET_RLM_CLint
ExpansionRx-OpenADMET RLM CLint
RLM CLint (rat liver microsomal intrinsic clearance) dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict the rat liver microsomal intrinsic clearance (RLM CLint) of molecules.
Note that this dataset was not part of the original challenge. It was provided by the organizers afterward as an additional endpoint.
Characteristic
Description
Tasks
1… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_RLM_CLint.forensic-fingerprint-dataset
Fingerprint Database
Dataset comprises 6,000+ fingerprint images from 100 individuals, with samples covering both hands and all ten fingers per person. It is designed for fingerprint identification, recognition systems, and forensic investigations, offering a robust resource for biometric data analysis, criminal investigations, and automated fingerprint matching.
By leveraging this dataset, researchers and law enforcement agencies can enhance fingerprint recognition algorithms… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/forensic-fingerprint-dataset.ExpansionRx_OpenADMET_KSOL
ExpansionRx-OpenADMET KSOL
KSOL dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict KSOL of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
7298
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1: ExpansionRx-OpenADMET Blind Challenge"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_KSOL.fingerprint-outputsMolPILE
MolPILE dataset
ArXiv preprint: "MolPILE - large-scale, diverse dataset for molecular representation learning" J. Adamczyk, J. Poziemski, F. Job, M. Król, M. Makowski
GitHub repository: https://github.com/scikit-fingerprints/MolPILE_dataset
Description
MolPILE is a large-scale molecular dataset designed for pretraining and evaluating machine learning models in cheminformatics.
It is compiled from several major chemical databases: UniChem, PubChem, Mcule, ChemSpace… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MolPILE.TDC_herg_central_at_10umtest
TDC_pampa_ncats
TDC PAMPA NCATS
PAMPA NCATS dataset [1], part of TDC [2] benchmark. It is intended to be used through
scikit-fingerprints library.
NCATS subset of PAMPA dataset.
PAMPA (parallel artificial membrane permeability assay) is an assay to evaluate drug permeability across the cellular membrane. The task models only the passive membrane diffusion. This is “NCATS” subset of the dataset created at National Center for Advancing Translational Sciences (NCATS).
This dataset is a part of… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/TDC_pampa_ncats.TDC_herg_central_at_1umTDC_caco2_wang
TDC Caco-2 Wang
Caco-2 Wang dataset [1], part of TDC [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict the rate at which drug passes through Caco-2 cells that serve as in vitro simulation of human intestinal tissue.
This dataset is a part of "absorption" subset of ADME tasks.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
910
Recommended splitscaffold
Recommended metric
MAE… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/TDC_caco2_wang.ExpansionRx_OpenADMET_MGMB
ExpansionRx-OpenADMET MGMB
MGMB dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict MGMB of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
431
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1: ExpansionRx-OpenADMET Blind Challenge"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_MGMB.TDC_solubility_aqsoldb
TDC Solubility AqSolDB
Solubility AqSolDB dataset dataset [1], part of TDC [2] benchmark. It is intended to be used through
scikit-fingerprints library.
he task is to predict the aqeuous solubility - a measure drug's ability to dissolve in water.
Poor water solubility could lead to slow drug absorptions,
inadequate bioavailablity and even induce toxicity.
This dataset is a part of "absorption" subset of ADME tasks.
Characteristic
Description
Tasks
1
Task type… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/TDC_solubility_aqsoldb.
