datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
selfies-ids-cleanedPubChem-124M-SMILES-SELFIES-InChI-IUPAC
PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC
Dataset Summary
This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format.
Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource:
SMILES: Raw and RDKit-Canonicalized.
SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.belka-selfies-idsbelka-selfies-train-cls-ftsmiles-selfies-pretrainPubChem-124M-SMILES-SELFIES-InChI-IUPAC
PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC
Dataset Summary
This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format.
Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource:
SMILES: Raw and RDKit-Canonicalized.
SELFIES: Pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/Bilsteen/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.selfies-train-idspubchem-selfies-pretrainchemnlp-selfies-pretrainselfie2animePubChem-124M-SMILES-SELFIES-InChI-IUPAC
PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC
Dataset Summary
This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format.
Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource:
SMILES: Raw and RDKit-Canonicalized.
SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/th-laurel/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.coconut-chembl34-selfies-mlm
Dataset Card for COCONUT+ChemBL34 SELFIES for MLM training (unmasked)
This dataset is a collection of molecular structures represented as SELFIES (Self-Referencing Embedded Strings), created by combining and processing data from COCONUTDB and ChemBL34. It contains 2,700,462 unique molecules across 13 chunks.
The dataset is specifically designed for pre-training language models on molecular representations using the Masked Language Model (MLM) approach. It consists of a single column… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/coconut-chembl34-selfies-mlm.PubChem10M_SELFIESPubChem10M dataset by DeepChem encoded to SELFIES using group-selfies.
belka-selfies-train-cls-ft-idsZINC-selfies-20mPubChem-SMILES-SELFIES-InChI-IUPAC-v2enamine-nature-selfies-pretrain-p1PubChem10M_SMILES_SELFIESpubchem-10m-selfiesChEMBL36-SELFIES
ChEMBL 36 SELFIES
Pre-training dataset for ModernMolBERT. Contains ~2.4M drug-like small molecules from ChEMBL 36 represented as SELFIES strings.
Dataset details
field
value
source
lukaskim/ChEMBL-36
representation
SELFIES
train rows
2,390,314
validation rows
24,228
total rows
2,414,542
min heavy atoms
3
max heavy atoms
100
max MW
1000.0
deduplicated by
InChIKey
split method
deterministic hash on InChIKey
valid fraction
0.01… See the full description on the dataset page: https://huggingface.co/datasets/HauserGroup/ChEMBL36-SELFIES.PubChem10M_SELFIES_TokenizedCustom cl100k tokenized version of PubChem10M_SELFIES.
PubChem-SMILES-SELFIES-InChI-IUPAC-v3selfiekyc_deepfake_detection
selfiekyc_deepfake_detection
Descripción General
Este repositorio contiene el código fuente, datasets procesados y experimentos desarrollados para la detección de imágenes faciales sintéticas o manipuladas mediante Inteligencia Artificial en escenarios de verificación de identidad corporativa tipo selfie-KYC.
El proyecto analiza la robustez de diferentes arquitecturas de Deep Learning frente a degradaciones operativas reales, incluyendo:
Compresión JPEG fuerte… See the full description on the dataset page: https://huggingface.co/datasets/juliopaciello/selfiekyc_deepfake_detection.enamine-diverse-selfies-pretrainSELFormer-selfiesftinder_selfies
Dataset Card for "ftinder_selfies"
More Information needed
molenet-selfies-pretrainPLINDER-selfiesguacamol-selfies-pretrainGLASS_GPCR_SELFIES
From the GLASS GPCR database, converted to SELFIES
Steps to prepare the database:
Download the GLASS database
wget https://zhanggroup.org/GLASS/downloads/interactions_active.tsv
wget https://zhanggroup.org/GLASS/downloads/interactions_inactives.tsv
wget https://zhanggroup.org/GLASS/downloads/targets.tsv
wget https://zhanggroup.org/GLASS/downloads/ligands.tsv
Select just the columns of interest
cut -d$'\t' -f6,9 ligands.tsv > ligands2.tsv
cut -d$'\t' -f2,5 targets.tsv >… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/GLASS_GPCR_SELFIES.
