datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
selfies-ids-cleanedPubChem-124M-SMILES-SELFIES-InChI-IUPAC
PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC
Dataset Summary
This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format.
Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource:
SMILES: Raw and RDKit-Canonicalized.
SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.Selfie-with-ID
Selfie Identity Dataset - 2 ID photo, 13 selfie
The dataset contains 65,000+ photo of more than 5,000 people from 40 countries, making it a valuable resource for exploring and developing identity verification solutions. This collection serves as a valuable resource for researchers and developers working on biometric verification solutions, especially in areas like facial recognition and financial services.
By utilizing this dataset, researchers can develop more robust… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/Selfie-with-ID.belka-selfies-idsbelka-selfies-train-cls-ftsmiles-selfies-pretrainPubChem-124M-SMILES-SELFIES-InChI-IUPAC
PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC
Dataset Summary
This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format.
Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource:
SMILES: Raw and RDKit-Canonicalized.
SELFIES: Pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/Bilsteen/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.selfies-train-idstea-app-selfies
Tea app selfies
This dataset contains selfies from the Tea app breach. Tea was an app that only allowed women to join and required users to submit both a selfie and a photograph of their ID before allowing entry.
This repository contains only the selfie images. ID photographs were filtered out before upload.
Contents
The dataset consists of approximately 6.7k JPEG files named with UUID-style filenames.
Intended use
This dataset may be useful for research into… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/tea-app-selfies.pubchem-selfies-pretrainkids-and-teens-selfie-dataset
Age Estimation
The dataset consists of 6,000 high-quality facial images from 300 people (children and teenagers), featuring a diverse range of facial features, poses, and attributes. It is designed for research and development in age estimation, facial recognition for younger demographics, and understanding media use patterns on social media platforms.
By utilizing this dataset, researchers and developers can advance their models for responsible technology, ensuring safer… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/kids-and-teens-selfie-dataset.chemnlp-selfies-pretrainselfies_and_id4083 sets, which includes 2 photos of a person from his documents and
13 selfies. 571 sets of Hispanics and 3512 sets of Caucasians.
Photo documents contains only a photo of a person.
All personal information from the document is hidden.selfie2animeselfie_and_video
Selfies and video dataset
4000 people in this dataset. Each person took a selfie on a webcam, took a selfie on a mobile phone. In addition, people recorded video from the phone and from the webcam, on which they pronounced a given set of numbers.
Includes folders corresponding to people in the dataset. Each folder includes 8 files (4 images and 4 videos).
💴 For Commercial Usage: To discuss your requirements, learn about the price and buy the dataset, leave a request on… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/selfie_and_video.Selfie_and_Official_ID_Photo_Dataset12,000+ people, 150,000+ images. Selfie with ID dataset for KYC verification, face identification and biometric training. Selfies paired with 2 official ID photos (passport, ID card, driver's license, residence permit). 10-15 photos per person with balanced demographics across ethnicity (Caucasian, Black, Asian, Latin American), gender and age (18-65).
Contact us and share your feedback - recieve additional samples for free! 😊
Key Highlights:
12,000+ real individuals… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/Selfie_and_Official_ID_Photo_Dataset.PubChem-124M-SMILES-SELFIES-InChI-IUPAC
PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC
Dataset Summary
This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format.
Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource:
SMILES: Raw and RDKit-Canonicalized.
SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/th-laurel/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.coconut-chembl34-selfies-mlm
Dataset Card for COCONUT+ChemBL34 SELFIES for MLM training (unmasked)
This dataset is a collection of molecular structures represented as SELFIES (Self-Referencing Embedded Strings), created by combining and processing data from COCONUTDB and ChemBL34. It contains 2,700,462 unique molecules across 13 chunks.
The dataset is specifically designed for pre-training language models on molecular representations using the Masked Language Model (MLM) approach. It consists of a single column… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/coconut-chembl34-selfies-mlm.PubChem10M_SELFIESPubChem10M dataset by DeepChem encoded to SELFIES using group-selfies.
male-selfie-image-dataset
Face Recognition, Face Detection, Male Photo Dataset 👨
The dataset is created on the basis of Selfies and ID Dataset
110,000+ photos of 74,000+ men from 141 countries.
The dataset includes photos of people's faces. All people presented in the dataset are men. The dataset contains a variety of images capturing individuals from diverse backgrounds and age groups.
Our dataset will diversify your data by adding more photos of men of different ages and ethnic groups… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/male-selfie-image-dataset.pubchem_selfiesThis dataset contains ~100M molecules from PubChem, with their SMILES and SELFIES representations.belka-selfies-train-cls-ft-idsZINC-selfies-20mPubChem-SMILES-SELFIES-InChI-IUPAC-v2enamine-nature-selfies-pretrain-p1selfie-and-video-datasetFace recognition dataset with 89,000+ selfies and videos from 5,600+ people. This dataset for face recognition is designed for training facial recognition models, face recognition training datasets, and facial recognition databases. Includes multi-ethnic subjects, varying lighting, pose variations, and video recordings per individual.
Contact us and share your feedback - recieve additional samples for free! 😊
Key Highlights:
5,600+ unique individuals across ages… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/selfie-and-video-dataset.female-selfie-image-dataset
Face Recognition, Face Detection, Female Photo Dataset 👩
The dataset is created on the basis of Selfies and ID Dataset
90,000+ photos of 46,000+ women from 141 countries.
The dataset includes photos of people's faces. All people presented in the dataset are women. The dataset contains a variety of images capturing individuals from diverse backgrounds and age groups.
Our dataset will diversify your data by adding more photos of women of different ages and ethnic groups… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/female-selfie-image-dataset.selfie-and-video-on-back-cameraThe dataset consists of selfies and video of real people made on a back camera
of the smartphone. The dataset solves tasks in the field of anti-spoofing and
it is useful for buisness and safety systems.PubChem10M_SMILES_SELFIESfencing-lerobot-realsense-depth-and-selfie-cam
fencing-lerobot-realsense-depth-and-selfie-cam
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
