datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PD12M
PD12M
Summary
At 12.4 million image-caption pairs, PD12M is the largest public domain image-text dataset to date, with sufficient size to train foundation models while minimizing copyright concerns. Through the Source.Plus platform, we also introduce novel, community-driven dataset governance mechanisms that reduce harm and support reproducibility over time.
Jordan Meyer Nicholas Padgett Cullen Miller Laura Exline
Paper Datasheet Project
About… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/PD12M.wine-reviews
Original Dataset Details
License: CC BY-NC-SA 4.0
Attribution: Zackthoutt
Source: Wine Reviews Dataset on Kaggle
CornellMovieDialogCorpusCornell Movie-Dialogs Corpus
Distributed together with:
"Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs"
Cristian Danescu-Niculescu-Mizil and Lillian Lee
Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, ACL 2011.
(this paper is included in this zip file)
NOTE: If you have results to report on these corpora, please send email to cristian@cs.cornell.edu or llee@cs.cornell.edu so we can add you to… See the full description on the dataset page: https://huggingface.co/datasets/spawn99/CornellMovieDialogCorpus.PD3M
PD3M
Summary
At 3.3 million image-caption pairs, PD3M is a subset of PD12M, containing images only with the highest aesthetic scores.
PD12M is the largest public domain image-text dataset to date, with sufficient size to train foundation models while minimizing copyright concerns. Through the Source.Plus platform, we also introduce novel, community-driven dataset governance mechanisms that reduce harm and support reproducibility over time.
Jordan Meyer Nicholas… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/PD3M.megalith-cc0
Megalith-CC0
A CC0-filtered version of the Megalith-10m dataset. The images have also been persisted to an independent public S3 bucket, supported by the AWS Open Data Registry program, for durability.
Why filter by CC0?
The images in Megalith-10m, having been gathered from Flickr, have attached licenses of CC0 and public domain. However, it is not clear if users assigning the public domain license to their works understand the implications of the public domain… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/megalith-cc0.GPQA-diamond-ClaudeR1
Dataset Card for GPQA Diamond Reasoning Benchmark
Dataset Details
Dataset Description
A benchmark dataset for evaluating hybrid AI architectures, comparing reasoning-augmented LLMs (DeepSeek R1) against standalone models (Claude Sonnet 3.5). Contains 198 physics questions with:
Ground truth answers and explanations
Model responses from multiple architectures
Granular token usage and cost metrics
Difficulty metadata and domain categorization
Curated by: LLM… See the full description on the dataset page: https://huggingface.co/datasets/spawn99/GPQA-diamond-ClaudeR1.PersuasionForGoodPersuasion for Good: Towards a Personalized Persuasive Dialogue System for Social Good
Dataset and Codebase for Persuasion for Good: Towards a Personalized Persuasive Dialogue System for Social Good
published as a long paper in ACL 2019.
https://arxiv.org/abs/1906.06725
If you use the datasets or any source codes included in this repository in your
work, please cite the following paper. The bibtex is listed below:
@article{wang2019persuasion,
title={Persuasion for Good: Towards a Personalized… See the full description on the dataset page: https://huggingface.co/datasets/spawn99/PersuasionForGood.ai-jailbreak
Dataset Card for ai-jailbreak
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/SpawnedShoyo/ai-jailbreak/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/SpawnedShoyo/ai-jailbreak.UK-Car-Plate-VRN-Dataset
Dataset Description
This dataset contains VRNs along with their respective UK car plate images. Each VRN record provides:
front_plate: Synthetically created front license plate image.
rear_plate: Synthetically created rear license plate image.
augmented_front_plate: Synthetically augmented front license plate image.
augmented_rear_plate: Synthetically augmented rear license plate image.
Dataset Creation Date
24 February 2025
Dataset Creation Method… See the full description on the dataset page: https://huggingface.co/datasets/spawn99/UK-Car-Plate-VRN-Dataset.gr00t-wiggle-diner-300-v2-spawnword-ladder-reasoningFranka_Apple_spawn_2_locationsFranka_Apple_spawn_2_locations_randomLightFranka_Apple_spawn_2_locations_EEemgena_os_process_spawn_tree_limiter_mcp_teaser
🚀 OS Security - Desktop Agent Process Tree & Fork-Bomb Limiter (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Scenarios + Executable MCP Server)🏆 Get the Full Production Package & Commercial EULA on Gumroad:👉 Purchase Full Package on Gumroad🏷️ Use coupon code LAUNCH20 for €20 off at checkout!
🌟 Domain Overview & Features
Tracks and bounds child process spawning tree depth to prevent runaway fork loops, zombie processes, and CPU hogging.… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_os_process_spawn_tree_limiter_mcp_teaser.
