datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MoleculeNet_PCBA
MoleculeNet PCBA
PCBA (PubChem BioAssay) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict biological activity against 128 bioassays, generated by high-throughput screening (HTS). All tasks are binary active/non-active.
Note that targets have missing values. Algorithms should be evaluated only on present labels. For training data, you may want to impute them, e.g. with zeros.
Characteristic… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_PCBA.MoleculeNet_Tox21
MoleculeNet Tox21
Tox21 dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict 12 toxicity targets, including nuclear receptors and stress response pathways. All tasks are binary.
Note that targets have missing values. Algorithms should be evaluated only on present labels. For training data, you may want to impute them, e.g. with zeros.
Characteristic
Description
Tasks
12
Task type
multitask… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_Tox21.LRGB_Peptides-func
LRGB Peptides-func
Peptides-func (Peptides functional) dataset, part of Long Range Graph Benchmark (LRGB) [1]. It is intended to be used through
scikit-fingerprints library.
The task is to predict functional properties of peptides.
Characteristic
Description
Tasks
10
Task type
classification
Total samples
15535
Recommended split
stratified random
Recommended metric
AUPRC
References
[1]
Dwivedi, Vijay Prakash, et al.
"Long Range Graph Benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/LRGB_Peptides-func.MoleculeNet_ClinTox
MoleculeNet ClinTox
Load and return the ClinTox dataset, part of MoleculeNet [1] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict drug approval viability, by predicting clinical trial toxicity and final FDA approval status. Both tasks are binary.
Characteristic
Description
Tasks
2
Task type
multitask classification
Total samples
1477
Recommended split
scaffold
Recommended metric
AUROC
References
[1]… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ClinTox.MoleculeNet_SIDER
MoleculeNet SIDER
Load and return the SIDER (Side Effect Resource) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict adverse drug reactions (ADRs) as drug side effects to 27 system organ classes in MedDRA classification. All tasks are binary.
Characteristic
Description
Tasks
12
Task type
multitask classification
Total samples
7831
Recommended split
scaffold
Recommended metric… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_SIDER.MoleculeNet_MUV
MoleculeNet MUV
Load and return the MUV (Maximum Unbiased Validation) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict 17 targets designed for validation of virtual screening techniques, based on PubChem BioAssays. All tasks are binary.
Note that targets have missing values. Algorithms should be evaluated only on present labels. For training data, you may want to impute them, e.g. with zeros.… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_MUV.MoleculeNet_ToxCast
MoleculeNet ToxCast
ToxCast dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict 617 toxicity targets from a large library of compounds based on in vitro high-throughput screening. All tasks are binary.
Note that targets have missing values. Algorithms should be evaluated only on present labels. For training data, you may want to impute them, e.g. with zeros.
Characteristic
Description
Tasks… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ToxCast.LRGB_Peptides-struct
LRGB Peptides-struct
Peptides-struct (Peptides structural) dataset, part of Long Range Graph Benchmark (LRGB) [1]. It is intended to be used through
scikit-fingerprints library.
The task is to predict structural properties of peptides. Note that this is raw data, whereas the original paper [1] specifies
that targets should be standardized (mean 0, standard deviation 1) before training and evaluation. scikit-fingerprints does this by
default in the loader function, otherwise this… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/LRGB_Peptides-struct.latent-inspector-fingerprints
latent-inspector fingerprints
Reference representation-geometry fingerprints for four self-supervised vision encoders — DINOv2 ViT-L/14, I-JEPA ViT-H/14, V-JEPA 2 ViT-L/16, and EUPE ViT-B/16 — computed on the same canonical image with latent-inspector.
This dataset is the numeric evidence layer behind the README table in abdelstark/vjepa2-vitl-fpc2-256-onnx. The ONNX exports in the Latent Inspector — ONNX Vision Encoders collection are the models; this dataset is what their patch… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/latent-inspector-fingerprints.
