datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
All-Prompt-JailbreakAI-Puppet-Theater-Actor-SFT
AI Puppet Theater Actor SFT
Synthetic supervised fine-tuning data for the Actor agent in AI Puppet Theater.
The dataset teaches a small language model to respond to a single puppet-theater beat with one compact JSON object. It is intended for hackathon prototyping, schema following, and local adapter experiments, not as a general storytelling or chat dataset.
Schema
Each row is chat-style JSONL:
{
"id": "actor-sft-v0-000001",
"source_mix": ["synthetic_v0"… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/AI-Puppet-Theater-Actor-SFT.compliance_eu_ai_act_bafin_dora_suite_teaser
🚀 Compliance & Governance - EU AI Act & BaFin/DORA Technical Compliance Suite (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (1,000 Samples) & Commercial EULA on Gumroad:👉 Compliance & Governance - EU AI Act & BaFin/DORA Technical Compliance Suite on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
📦 What is Inside the Full Production Package:
1,000 Verified FAANG v2.0… See the full description on the dataset page: https://huggingface.co/datasets/emgena/compliance_eu_ai_act_bafin_dora_suite_teaser.ToxicDataset
Comprehensive Toxic Content Dataset
Dataset Description
This dataset contains 1,000,000 synthetically generated records of toxic, abusive, harmful, and offensive content designed for training content moderation systems and hate speech detection models.
Dataset Summary
This comprehensive dataset includes multiple categories of toxic content:
Toxic content (insults, derogatory terms)
Abusive language patterns
Gender bias statements
Dangerous/threatening content… See the full description on the dataset page: https://huggingface.co/datasets/AiActivity/ToxicDataset.legal-ai-act-spanish-sft-7k⚠️ Legal and Liability Disclaimer
This dataset is provided for research and educational purposes only.
It does not constitute legal advice, nor does it represent an official or authoritative interpretation of Regulation (EU) 2024/1689 (EU AI Act).
The content is synthetically generated and may contain errors, omissions, or hallucinations.
Under no circumstances should this dataset be used as a basis for legal, compliance, or regulatory decision-making.
The authors disclaim any liability for… See the full description on the dataset page: https://huggingface.co/datasets/hugoramallo/legal-ai-act-spanish-sft-7k.governed-ai-actions-bench
Governed AI Actions Bench
This dataset evaluates whether a governed AI system can route requests into
policy actions: allow, refuse, rewrite, summarize, escalate, and shadow-mode
detect-but-allow. Each row contains an endpoint-agnostic prompt, policy config,
expected decision metadata, and behavioral checks.
The benchmark is designed for teams that need more than binary moderation. A
runner can use the rows to verify policy metadata, reason codes, rollout mode,
final-content… See the full description on the dataset page: https://huggingface.co/datasets/abliterationaiorg/governed-ai-actions-bench.customer-feedback-action-plans
Customer Feedback → Action Plans
A small, practical dataset that maps raw customer feedback (e.g., restaurant reviews) to actionable recommendations with optional aspect annotations and reasoning. Useful for training instruction-following models, aspect-aware summarizers, or classification heads that support the generation task.
Files & Splits
train.csv — main training split for generation.
validation.csv — validation split for generation.
train_aux_classification.csv —… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/customer-feedback-action-plans.governed-ai-actions-bench
Governed AI Actions Bench
This dataset evaluates whether a governed AI system can route requests into
policy actions: allow, refuse, rewrite, summarize, escalate, and shadow-mode
detect-but-allow. Each row contains an endpoint-agnostic prompt, policy config,
expected decision metadata, and behavioral checks.
The benchmark is designed for teams that need more than binary moderation. A
runner can use the rows to verify policy metadata, reason codes, rollout mode,
final-content… See the full description on the dataset page: https://huggingface.co/datasets/abliterationai/governed-ai-actions-bench.
