chembl
Datasets
All datasets matching “chembl”ChEMBLChEMBL-36
ChEMBL 36
ChEMBL 36 converted to HuggingFace datasets format. Source data is from the ChEMBL database (EMBL-EBI), a manually curated database of bioactive molecules with drug-like properties.
Dataset configs
molecules (~2.4M rows)
All compounds in ChEMBL with canonical SMILES representations.
Column
Type
Description
chembl_id
string
ChEMBL compound identifier
canonical_smiles
string
Canonical SMILES representation
standard_inchi… See the full description on the dataset page: https://huggingface.co/datasets/lukaskim/ChEMBL-36.smiles-molecules-chembl
ChEMBL Molecule Generation Dataset
Dataset Description
ChEMBL is a manually curated database of bioactive molecules with drug-like properties. It brings together chemical, bioactivity and genomic data to aid the translation of genomic information into effective new drugs.
Task Description
For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties measured by some oracles.… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-chembl.chembl-smiles-curated
Dataset Card for Curated ChEMBL SMILES Dataset (lovingscience/chembl-smiles-curated)
This dataset is a reproducibly cleaned, deduplicated, and split ChEMBL SMILES dataset built for training molecular generative models and chemical space mapping tools (such as Imolgen and Navifold).
Dataset Pipeline Provenance & Configuration
ChEMBL Release Version: 37
RDKit Version: 2026.03.6
Deduplication & Split Seed: 42
Split Ratios: Train (80.0%) / Validation (10.0%) / Test… See the full description on the dataset page: https://huggingface.co/datasets/lovingscience/chembl-smiles-curated.MoleculeACE_chembl214_ki
MoleculeACE ChEMBL214 Ki
ChEMBL214 dataset, originally part of ChEMBL database [1], processed in MoleculeACE [2] for activity cliff evaluation. It is intended to be use through scikit-fingerprints library.
The task is to predict the inhibitor constant (Ki) of molecules against the 5-hydroxytryptamine receptor 1a target.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
3317
Recommended splitactivity_cliff
Recommended metric
RMSE… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeACE_chembl214_ki.MoleculeACE_chembl204_ki
MoleculeACE ChEMBL204 Ki
ChEMBL204 dataset, originally part of ChEMBL database [1], processed in MoleculeACE [2] for activity cliff evaluation. It is intended to be use through scikit-fingerprints library.
The task is to predict the inhibitor constant (Ki) of molecules against the Prothrombin target.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
2754
Recommended split
activity_cliff
Recommended metric
RMSE
References
[1]
B.… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeACE_chembl204_ki.
