CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Spawning /pd12m-fullThis dataset is the downloaded variant of Spawning/PD12M. More specifically, this dataset is compatible with webdataset. It was made public after obtaining permission from the original authors of the dataset. You can use the following to explore the dataset with webdataset: import webdataset as wds dataset_path = "pipe:curl -s -f -L https://huggingface.co/datasets/sayakpaul/pd12m-full/resolve/main/{00155..02480}.tar" dataset = ( wds.WebDataset(dataset_path… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/pd12m-full.image10M<n<100M21 likes12k downloads2y agoHugging Face02Spawning /PD12M PD12M Summary At 12.4 million image-caption pairs, PD12M is the largest public domain image-text dataset to date, with sufficient size to train foundation models while minimizing copyright concerns. Through the Source.Plus platform, we also introduce novel, community-driven dataset governance mechanisms that reduce harm and support reproducibility over time. Jordan Meyer Nicholas Padgett Cullen Miller Laura Exline Paper Datasheet Project About… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/PD12M.image10M<n<100M186 likes3.4k downloads2y agoHugging Face03Intelligent-Internet /pd12m PD12M This is a curated PD12M dataset for use with the II-Commons project. Dataset Details Dataset Description This dataset comprises a curated Public Domain 12M image collection, refined by filtering for active image links. EXIF data was extracted, and images underwent preprocessing and feature extraction using SigLIP 2. All vector embeddings are normalized 16-bit half-precision vectors optimized for L2 indexing with vectorchord.… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/pd12m.imagefeature-extraction10M<n<100M8 likes2.4k downloads1y agoHugging Face04mletrasdl /pd12m-subsetimage100K<n<1M0 likes200 downloads9mo agoHugging Face05umarigan /PD12M-TurkishTranslated from English to Tuskish language from: https://huggingface.co/datasets/Spawning/PD12M One of the biggest text-to-image dataset in Turkish language Metadata The metadata is made available through a series of parquet files with the following schema: text: Translated caption for the image. id: A unique identifier for the image. url: The URL of the image. caption: A caption for the image. width: The width of the image in pixels. height: The height of the image in pixels.… See the full description on the dataset page: https://huggingface.co/datasets/umarigan/PD12M-Turkish.imagequestion-answering10M<n<100M8 likes146 downloads2y agoHugging Face06Rinovative /pd12m_dct_based_synthetic_stegano PD12M DCT-Based Synthetic Steganography Dataset This dataset is a synthetically generated steganographic image dataset based on the PD12M (Public Domain 12M) image collection.It simulates detectable modifications produced by real-world JPEG steganography algorithms, using only public domain data. Each original image is duplicated into three synthetic stego variants, inspired by real-world JPEG steganographic algorithms: synthetic_JMiPOD: simulated using conseal.nsF5, as JMiPOD is… See the full description on the dataset page: https://huggingface.co/datasets/Rinovative/pd12m_dct_based_synthetic_stegano.image1K<n<10K0 likes80 downloads1y agoHugging Face07umarigan /PD12M-Turkish-Images-cleaned_37kimage10K<n<100K2 likes28 downloads1y agoHugging Face08d0rj /PD12M-ruTranslated captions from Spawning/PD12M into Russian using Google Translate. imageimage-feature-extraction10M<n<100M2 likes27 downloads2y agoHugging Face09umarigan /PD12M-Turkish-Imagesimage10K<n<100K0 likes13 downloads1y agoHugging Face10BootsofLagrangian /pd12m-recap-qwen3p5-35b-a3b PD12M recaptions with Qwen3.5-35B-A3B Dataset pd12m: 12.316 Million primary train rows plus 8.141 Million bundled extension rows. The default train split contains exactly the 12,316,334 successfully generated recaptions from the pinned full HF image surface. A user who loads only train receives one deterministic primary recap row per retained HF-full asset and does not need any local materialization. The bundle split contains 8,140,780 additional captions from the UUID-verified… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/pd12m-recap-qwen3p5-35b-a3b.image10M<n<100M0 likes12 downloads1mo agoHugging Face11ShuhongZheng /pd12mimage0 likes11 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.