datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smiles-molecules-chembl
ChEMBL Molecule Generation Dataset
Dataset Description
ChEMBL is a manually curated database of bioactive molecules with drug-like properties. It brings together chemical, bioactivity and genomic data to aid the translation of genomic information into effective new drugs.
Task Description
For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties measured by some oracles.… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-chembl.smishing-syntheticMOZ-Smishing
Dataset Summary
MOZ-Smishing is a benchmark dataset specifically designed for detecting smishing attacks targeting Mobile Money Transfer (MMT) systems. This dataset addresses the critical lack of publicly available SMS phishing datasets in this domain, especially for non-English languages. It comprises crowd-sourced text messages from Mozambican mobile users, meticulously annotated into two categories: legitimate messages and smishing attempts. The messages are primarily in… See the full description on the dataset page: https://huggingface.co/datasets/MOZNLP/MOZ-Smishing.ta-ESM2
taxonomy_aware_ESM2
This repository implements a Taxonomy-Aware Protein Function Prediction model. It synergizes the structural language understanding of ESM2 (Evolutionary Scale Modeling) with explicit phylogenetic lineage information.
text_emotionsmiles-molecules-moses
MOSES Molecule Generation Dataset
Dataset Description
Molecular Sets (MOSES) is a benchmark platform for distribution learning based molecule generation. Within this benchmark, MOSES provides a cleaned dataset of molecules that are ideal of optimization. It is processed from the ZINC Clean Leads dataset.
Task Description
For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-moses.camie-tagger-vs-wd-tagger-val
What's what
1_json_to_csv.py:
converts cm_tags.json to a csv format I'm more used to and that is easier to use with my existing tooling
2_retrieve_images_by_cc.py:
retrieves the validation images using cheesechaser, downloads them to "original/"
3_common_tags.py:
clean up both models tag sets to only consider the common tags; note down the indexes to use to fetch the correct tag probs from the dumps generated by the inference scripts
4_cm_onnx_inference.py:
run… See the full description on the dataset page: https://huggingface.co/datasets/SmilingWolf/camie-tagger-vs-wd-tagger-val.present-corner-5c5ec7
present-corner-5c5ec7
Synthetic weather test data: 45 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/steven-smith/present-corner-5c5ec7.worth-church-eaf4de
worth-church-eaf4de
Synthetic products test data: 50 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/michael-smith/worth-church-eaf4de.CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_conciseThe CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_concise dataset is a part of the study "Predicting Nanoparticle Effects on Small Biomolecule Functionalities Using the Capability of Scikit-learn and PyTorch: A Case Study on Inhibitors of the DNA Damage-Inducible Transcript 3 (CHOP)"
https://doi.org/10.48550/arXiv.2504.09537
Dataset content: The Dataset has 20,309 rows of samples, 12 columns of features. Each feature corresponds to a range of the 1H NMR spectroscopy chemical shifts scale. Natural… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_concise.SELFormer-smilesmalaria_CIDs_SIDs_SMILES_targets
Dataset was extracted from a dataset provided by NOVARTIS: Inhibition of Plasmodium falciparum W2 (drug-resistant) proliferation in erythrocyte-based infection assay https://pubchem.ncbi.nlm.nih.gov/bioassay/449704
Code of NNs https://pubchem.ncbi.nlm.nih.gov/bioassay/449704 utilising the resulting dataset.
smishing-AZ-SC
SMS Classification Dataset - Azerbaijan
Citation
If you use this dataset in your research, please cite:
@article{shahbazov2026azerbaijani,
title = {SMS dataset for multi-class classification of ham, spam, and smishing in Azerbaijani language},
author = {Vusal Shahbazov},
journal = {Problems of Information Technology},
volume = {17},
number = {1},
pages = {32--39},
year = {2026},
issn = {2304-0157},
doi =… See the full description on the dataset page: https://huggingface.co/datasets/VusalShahbazovAz/smishing-AZ-SC.organic_reactionsCHOP_inhibitors_SMILES_1H_NMR_spectrosopy_extensiveThe CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_extensive dataset is a part of the study "Predicting Nanoparticle Effects on Small Biomolecule Functionalities Using the Capability of Scikit-learn and PyTorch: A Case Study on Inhibitors of the DNA Damage-Inducible Transcript 3 (CHOP)"
https://doi.org/10.48550/arXiv.2504.09537
Dataset content: The Dataset has 20,309 rows of samples, 103 columns of features. Each feature corresponds to a range of the 1H NMR spectroscopy chemical shifts scale.… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_extensive.norton-proclamations-01CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_conciseThe CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_concise dataset is a part of the study "Predicting Nanoparticle Effects on Small Biomolecule Functionalities Using the Capability of Scikit-learn and PyTorch: A Case Study on Inhibitors of the DNA Damage-Inducible Transcript 3 (CHOP)"
https://doi.org/10.48550/arXiv.2504.09537
Dataset content: The Dataset has 20,309 rows of samples, 220 columns of features. Each feature corresponds to a range of the 13C NMR spectroscopy chemical shifts scale.… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_concise.bioactives-naturals-smiles-molgen
Valid Bioactives and Natural Product SMILES
~2.7M valid SMILES built and curated from ChemBL34 (Zdrazil et al. 2023), COCONUTDB (Sorokina et al. 2021), and Supernatural3 (Gallo et al. 2023) dataset.
Curated by: gbyuvd
References
BibTeX
COCONUTDB
@article{sorokina2021coconut,
title={COCONUT online: Collection of Open Natural Products database},
author={Sorokina, Maria and Merseburger, Peter and Rajan, Kohulan and Yirik, Mehmet Aziz and Steinbeck… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/bioactives-naturals-smiles-molgen.CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_extensiveThe CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_extensive dataset is a part of the study "Predicting Nanoparticle Effects on Small Biomolecule Functionalities Using the Capability of Scikit-learn and PyTorch: A Case Study on Inhibitors of the DNA Damage-Inducible Transcript 3 (CHOP)"
https://doi.org/10.48550/arXiv.2504.09537
Dataset content:
The Dataset has 20,309 rows of samples, 1,928 columns of features. Each feature corresponds to a range of the 13C NMR spectroscopy chemical shifts scale.… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_extensive.human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_with_molecular! A note: To address the size limitations on Hugging Face, only 200 of the 59,609 rows were uploaded. The full dataset is available upon request for interested parties
The human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_with_molecular_features dataset is a part of the study " Comparative analysis of computational approaches for predicting Transthyretin (TTR) transcription activators and human dopamine D1 receptor antagonists"
https://doi.org/10.48550/arXiv.2506.01137
The… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_with_molecular.runmke-novel-druglike-smiles
MKE Novel Drug-Like Molecules — AI-Generated SMILES Dataset
Commercial dataset available for purchase. Get access on our website →
Overview
This dataset contains 9,274 AI-generated, novel, drug-like small molecules, rigorously validated and filtered for pharmaceutical relevance. Every molecule in this dataset is:
✅ Chemically valid — RDKit-verified SMILES
✅ 100% novel — verified against 4,643,595 known compounds from MOSES, ZINC-250k, and ChEMBL
✅ Drug-like — QED… See the full description on the dataset page: https://huggingface.co/datasets/MKEChem/mke-novel-druglike-smiles.newsKMAC_Testsmit_resumeMOZ-Smishing
Dataset Summary
MOZ-Smishing is a benchmark dataset specifically designed for detecting smishing attacks targeting Mobile Money Transfer (MMT) systems. This dataset addresses the critical lack of publicly available SMS phishing datasets in this domain, especially for non-English languages. It comprises crowd-sourced text messages from Mozambican mobile users, meticulously annotated into two categories: legitimate messages and smishing attempts. The messages are primarily in… See the full description on the dataset page: https://huggingface.co/datasets/chisomobanja/MOZ-Smishing.vansh_dataTestSMILES_Big_Data_SetSYN-RoadGps
SYN-RoadGPS
SYN-RoadGPS is a road network-constrained synthetic GPS trajectory dataset for trajectory modeling and evaluation.
The dataset contains synthetic trajectories represented by road segment identifiers, timestamps, within-segment position ratios, and reconstructed GPS coordinates.
The raw source GPS trajectory records are not publicly released due to privacy and access restrictions. The released data are synthetic and are intended for research on synthetic trajectory… See the full description on the dataset page: https://huggingface.co/datasets/Smile1120/SYN-RoadGps.
