CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HoangHa /selfies-ids-cleanedtext100M<n<1B0 likes739 downloads2y agoHugging Face02hheiden /PubChem-124M-SMILES-SELFIES-InChI-IUPAC PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format. Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource: SMILES: Raw and RDKit-Canonicalized. SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.texttext-generation100M<n<1B8 likes689 downloads9mo agoHugging Face03HoangHa /belka-selfies-idstext10M<n<100M0 likes484 downloads2y agoHugging Face04HoangHa /belka-selfies-train-cls-fttabular10M<n<100M0 likes466 downloads2y agoHugging Face05HoangHa /smiles-selfies-pretraintext100M<n<1B3 likes462 downloads2y agoHugging Face06Bilsteen /PubChem-124M-SMILES-SELFIES-InChI-IUPAC PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format. Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource: SMILES: Raw and RDKit-Canonicalized. SELFIES: Pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/Bilsteen/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.texttext-generation100M<n<1B0 likes436 downloads2mo agoHugging Face07HoangHa /selfies-train-idstext100M<n<1B0 likes418 downloads2y agoHugging Face08HoangHa /pubchem-selfies-pretraintext100M<n<1B1 likes344 downloads2y agoHugging Face09HoangHa /chemnlp-selfies-pretraintext10M<n<100M0 likes219 downloads2y agoHugging Face10huggan /selfie2animetext1K<n<10K3 likes186 downloads4y agoHugging Face11th-laurel /PubChem-124M-SMILES-SELFIES-InChI-IUPAC PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format. Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource: SMILES: Raw and RDKit-Canonicalized. SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/th-laurel/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.texttext-generation100M<n<1B0 likes148 downloads6mo agoHugging Face12gbyuvd /coconut-chembl34-selfies-mlm Dataset Card for COCONUT+ChemBL34 SELFIES for MLM training (unmasked) This dataset is a collection of molecular structures represented as SELFIES (Self-Referencing Embedded Strings), created by combining and processing data from COCONUTDB and ChemBL34. It contains 2,700,462 unique molecules across 13 chunks. The dataset is specifically designed for pre-training language models on molecular representations using the Masked Language Model (MLM) approach. It consists of a single column… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/coconut-chembl34-selfies-mlm.text100K<n<1M0 likes145 downloads1y agoHugging Face13alxfgh /PubChem10M_SELFIESPubChem10M dataset by DeepChem encoded to SELFIES using group-selfies. text1M<n<10M1 likes143 downloads3y agoHugging Face14HoangHa /belka-selfies-train-cls-ft-idstabular10M<n<100M0 likes77 downloads2y agoHugging Face15victornica /ZINC-selfies-20mtext10M<n<100M0 likes66 downloads1y agoHugging Face16altaidevorg /PubChem-SMILES-SELFIES-InChI-IUPAC-v2image100K<n<1M0 likes65 downloads8mo agoHugging Face17HoangHa /enamine-nature-selfies-pretrain-p1text100M<n<1B0 likes60 downloads2y agoHugging Face18mikemayuare /PubChem10M_SMILES_SELFIEStext10M<n<100M3 likes47 downloads2y agoHugging Face19Neeze /pubchem-10m-selfiestext1M<n<10M0 likes34 downloads10mo agoHugging Face20HauserGroup /ChEMBL36-SELFIES ChEMBL 36 SELFIES Pre-training dataset for ModernMolBERT. Contains ~2.4M drug-like small molecules from ChEMBL 36 represented as SELFIES strings. Dataset details field value source lukaskim/ChEMBL-36 representation SELFIES train rows 2,390,314 validation rows 24,228 total rows 2,414,542 min heavy atoms 3 max heavy atoms 100 max MW 1000.0 deduplicated by InChIKey split method deterministic hash on InChIKey valid fraction 0.01… See the full description on the dataset page: https://huggingface.co/datasets/HauserGroup/ChEMBL36-SELFIES.tabularfill-mask1M<n<10M0 likes32 downloads4mo agoHugging Face21alxfgh /PubChem10M_SELFIES_TokenizedCustom cl100k tokenized version of PubChem10M_SELFIES. text1M<n<10M2 likes31 downloads3y agoHugging Face22altaidevorg /PubChem-SMILES-SELFIES-InChI-IUPAC-v3image100K<n<1M0 likes31 downloads8mo agoHugging Face23juliopaciello /selfiekyc_deepfake_detection selfiekyc_deepfake_detection Descripción General Este repositorio contiene el código fuente, datasets procesados y experimentos desarrollados para la detección de imágenes faciales sintéticas o manipuladas mediante Inteligencia Artificial en escenarios de verificación de identidad corporativa tipo selfie-KYC. El proyecto analiza la robustez de diferentes arquitecturas de Deep Learning frente a degradaciones operativas reales, incluyendo: Compresión JPEG fuerte… See the full description on the dataset page: https://huggingface.co/datasets/juliopaciello/selfiekyc_deepfake_detection.imageimage-classification1K<n<10K0 likes31 downloads2mo agoHugging Face24HoangHa /enamine-diverse-selfies-pretraintext10M<n<100M0 likes21 downloads2y agoHugging Face25HUBioDataLab /SELFormer-selfiestext1M<n<10M0 likes18 downloads3y agoHugging Face26emilianJR /ftinder_selfies Dataset Card for "ftinder_selfies" More Information needed imagen<1K1 likes17 downloads3y agoHugging Face27HoangHa /molenet-selfies-pretraintext100K<n<1M0 likes16 downloads2y agoHugging Face28Synthyra /PLINDER-selfiestabular10K<n<100K0 likes16 downloads8mo agoHugging Face29HoangHa /guacamol-selfies-pretraintext1M<n<10M0 likes15 downloads2y agoHugging Face30introvoyz041 /GLASS_GPCR_SELFIES From the GLASS GPCR database, converted to SELFIES Steps to prepare the database: Download the GLASS database wget https://zhanggroup.org/GLASS/downloads/interactions_active.tsv wget https://zhanggroup.org/GLASS/downloads/interactions_inactives.tsv wget https://zhanggroup.org/GLASS/downloads/targets.tsv wget https://zhanggroup.org/GLASS/downloads/ligands.tsv Select just the columns of interest cut -d$'\t' -f6,9 ligands.tsv > ligands2.tsv cut -d$'\t' -f2,5 targets.tsv >… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/GLASS_GPCR_SELFIES.tabular100K<n<1M0 likes14 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.