spawn
Datasets
All datasets matching “spawn”pd12m-fullThis dataset is the downloaded variant of Spawning/PD12M. More specifically, this dataset
is compatible with webdataset. It was made public after obtaining permission
from the original authors of the dataset.
You can use the following to explore the dataset with webdataset:
import webdataset as wds
dataset_path = "pipe:curl -s -f -L https://huggingface.co/datasets/sayakpaul/pd12m-full/resolve/main/{00155..02480}.tar"
dataset = (
wds.WebDataset(dataset_path… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/pd12m-full.PD12M
PD12M
Summary
At 12.4 million image-caption pairs, PD12M is the largest public domain image-text dataset to date, with sufficient size to train foundation models while minimizing copyright concerns. Through the Source.Plus platform, we also introduce novel, community-driven dataset governance mechanisms that reduce harm and support reproducibility over time.
Jordan Meyer Nicholas Padgett Cullen Miller Laura Exline
Paper Datasheet Project
About… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/PD12M.Poketwo-Spawn-Imagesloracailin020wine-reviews
Original Dataset Details
License: CC BY-NC-SA 4.0
Attribution: Zackthoutt
Source: Wine Reviews Dataset on Kaggle
pd-extended
pd-extended
Summary
PD-Extended is a collection of ~34.7 million image/caption pairs derived from the PD12M and Megalith-CC0 datasets. The image/caption pairs are accompanied with metadata, such as mime type and dimensions, as well as the accompanying CLIP-L14 embeddings. Of note, these images retain their original licensing, and the source_id is available to pair any derived image to its source within the original dataset. All images are paired with synthetic captions… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/pd-extended.
