datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pd12m-fullThis dataset is the downloaded variant of Spawning/PD12M. More specifically, this dataset
is compatible with webdataset. It was made public after obtaining permission
from the original authors of the dataset.
You can use the following to explore the dataset with webdataset:
import webdataset as wds
dataset_path = "pipe:curl -s -f -L https://huggingface.co/datasets/sayakpaul/pd12m-full/resolve/main/{00155..02480}.tar"
dataset = (
wds.WebDataset(dataset_path… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/pd12m-full.PD12M
PD12M
Summary
At 12.4 million image-caption pairs, PD12M is the largest public domain image-text dataset to date, with sufficient size to train foundation models while minimizing copyright concerns. Through the Source.Plus platform, we also introduce novel, community-driven dataset governance mechanisms that reduce harm and support reproducibility over time.
Jordan Meyer Nicholas Padgett Cullen Miller Laura Exline
Paper Datasheet Project
About… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/PD12M.preprocessed_DCAE-f64_1024_pd12m-fullpd12m
PD12M
This is a curated PD12M dataset for use with the II-Commons project.
Dataset Details
Dataset Description
This dataset comprises a curated Public Domain 12M image collection, refined by filtering for active image links. EXIF data was extracted, and images underwent preprocessing and feature extraction using SigLIP 2. All vector embeddings are normalized 16-bit half-precision vectors optimized for L2 indexing with vectorchord.… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/pd12m.PD12M-256px_dc-ae-f32c32-sana-1.0pd12m-subsetPD12M-TurkishTranslated from English to Tuskish language from: https://huggingface.co/datasets/Spawning/PD12M
One of the biggest text-to-image dataset in Turkish language
Metadata
The metadata is made available through a series of parquet files with the following schema:
text: Translated caption for the image.
id: A unique identifier for the image.
url: The URL of the image.
caption: A caption for the image.
width: The width of the image in pixels.
height: The height of the image in pixels.… See the full description on the dataset page: https://huggingface.co/datasets/umarigan/PD12M-Turkish.preprocessed_DCAE-f64_pd12m-fullpd12m_dct_based_synthetic_stegano
PD12M DCT-Based Synthetic Steganography Dataset
This dataset is a synthetically generated steganographic image dataset based on the PD12M (Public Domain 12M) image collection.It simulates detectable modifications produced by real-world JPEG steganography algorithms, using only public domain data.
Each original image is duplicated into three synthetic stego variants, inspired by real-world JPEG steganographic algorithms:
synthetic_JMiPOD: simulated using conseal.nsF5, as JMiPOD is… See the full description on the dataset page: https://huggingface.co/datasets/Rinovative/pd12m_dct_based_synthetic_stegano.PD12M-splits-256px_dc-ae-f32c32-sana-1.0PD12M-Turkish-Images-cleaned_37kPD12M-ruTranslated captions from Spawning/PD12M into Russian using Google Translate.
PD12M-Turkish-Imagespd12m-recap-qwen3p5-35b-a3b
PD12M recaptions with Qwen3.5-35B-A3B
Dataset pd12m: 12.316 Million primary train rows plus 8.141 Million bundled extension rows.
The default train split contains exactly the 12,316,334 successfully generated recaptions from the pinned full HF image surface. A user who loads only train receives one deterministic primary recap row per retained HF-full asset and does not need any local materialization. The bundle split contains 8,140,780 additional captions from the UUID-verified… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/pd12m-recap-qwen3p5-35b-a3b.pd12m
