datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wildchat_filtered_english_description_label_no_moderationText-Moderation-Multilingual
Text-Moderation-Multilingual
A comprehensive multilingual text moderation dataset combining multiple high-quality sources for training robust content moderation classifiers.
Dataset Summary
This dataset aggregates text moderation data from multiple sources to create a large-scale, diverse training corpus for content moderation systems. It includes text samples labeled across multiple harmful content categories, supporting both multilingual and English-specific moderation… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/Text-Moderation-Multilingual.safety-moderation-benchmark
Safety Moderation Benchmark
Overview
A comprehensive benchmark for training binary safety classifiers to detect harmful content across 9 safety-critical policy domains. The dataset combines 100% synthetic evaluation data with diverse real-world and synthetic training samples.
Total Size: 228,925 samples
Train: 191,186 (83.5%)
Validation: 33,739 (14.7%)
Test: 4,000 (1.7% - 100% synthetic, stratified)
Dataset Composition
Sources… See the full description on the dataset page: https://huggingface.co/datasets/mvrcii/safety-moderation-benchmark.multi-lingual-prompt-moderation
Text-Moderation-Multilingual
A comprehensive multilingual text moderation dataset combining multiple high-quality sources for training robust content moderation classifiers.
Dataset Summary
This dataset aggregates text moderation data from multiple sources to create a large-scale, diverse training corpus for content moderation systems. It includes text samples labeled across multiple harmful content categories, supporting both multilingual and English-specific moderation… See the full description on the dataset page: https://huggingface.co/datasets/enguard/multi-lingual-prompt-moderation.Cockatoo-Moderation-V1
Cockatoo Moderation V1
This dataset is created by merging ucberkeley-dlab/measuring-hate-speech, KoalaAI/Text-Moderation-Multilingual, and google/civil_comments. Thus, this set is subject to different licenses (see below).
Size: 3,399,535 rows
Files:
Cockatoo-Moderation-V1_corpus.parquet: The unlabeled dataset
Cockatoo-Moderation-V1_labeled.parquet: Labeled (coming soon)
Method:
This dataset is primarily synthetically labeled with a small… See the full description on the dataset page: https://huggingface.co/datasets/DominicTWHV/Cockatoo-Moderation-V1.moderation-bias-benchmark
Moderation Bias: LLM Content Moderation Benchmark
One row per model evaluation of one prompt. This dataset is the raw audit log
behind moderationbias.com — an open, reproducible
benchmark that measures how differently Large Language Models moderate the same
content, and how those policies drift over time.
Homepage: https://moderationbias.com
Repository: https://github.com/jacobkandel/llm-content-moderation-analysis
Leaderboard: https://moderationbias.com/leaderboard
Paper /… See the full description on the dataset page: https://huggingface.co/datasets/jmk9494/moderation-bias-benchmark.merged_content_moderation_and_prompt_injection_newContent-Moderation-and-Safety
🇰🇿 Content Moderation and Safety, Kazakh Context
Dataset Summary
Content Moderation and Safety (Profanity) Kazakh Context is a comprehensive dataset designed specifically to train Large Language Models (LLMs) in detecting, classifying, and mitigating toxic, aggressive, or unsafe text in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
17,827
Total Words (approx.)
1,674,638
Avg.… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content-Moderation-and-Safety.moderation-test-resultsContent_Moderation_and_Safety_Kazakh_Context
🇰🇿 Content Moderation and Safety Kazakh Context
Dataset Summary
Toxic Speech Analysis and Mitigation, Kazakh Context is an advanced AI Safety dataset designed to train Large Language Models (LLMs) to detect, deeply analyze, and constructively rewrite toxic or harmful speech in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
12,063
Total Words (approx.)
5,869,718
Avg. Words per… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content_Moderation_and_Safety_Kazakh_Context.open-australian-legal-qa-paraphrased-moderation-resultsreddit-moderationenwiki-image-content-moderationThis dataset is composed of scores of images taken from English Wikipedia and Wikimedia Commons. The scores are the outputs of the models
https://github.com/bumble-tech/private-detector
https://huggingface.co/Freepik/nsfw_image_detector
https://huggingface.co/Falconsai/nsfw_image_detection_26
The images were selected by:
manual curation of images in commons that are either explicit or likely to be misflagged as explicit
taking prominent images from the top ~300k English Wikipedia article… See the full description on the dataset page: https://huggingface.co/datasets/derenrich/enwiki-image-content-moderation.qa_moderationopenai-moderation
Dataset
Homepage: https://github.com/openai/moderation-api-release
Description: A Holistic Approach to Undesired Content Detection
Citation:
@article{openai2022moderation,
title={A Holistic Approach to Undesired Content Detection},
author={Todor Markov and Chong Zhang and Sandhini Agarwal and Tyna Eloundou and Teddy Lee and Steven Adler and Angela Jiang and Lilian Weng},
journal={arXiv preprint arXiv:2208.03274},
year={2022}
}
t5-moderation-datasetmoderation_dataset
Moderation Dataset
Based on mmathys/openai-moderation-api-evaluation and davanstrien/WELFake
Warning
This dataset contains nsfw, chocking, discriminatory and hateful text.
It is intended to be used to train moderation AI assistants and should not be used for any other mean or reason.
Please use with care.
Category
Label
Definition
sexual
S
Content meant to arouse sexual excitement, such as the description of sexual activity, or that promotes sexual services (excluding… See the full description on the dataset page: https://huggingface.co/datasets/DaijobuAI/moderation_dataset.spanish-safety-moderation-guardrails-updatedmerged_content_moderation_and_prompt_injectionllm-moderationcontent-moderationopenai-moderation-binary
🧠 OpenAI Moderation Binary Dataset
This dataset is a binary-labeled version of the original OpenAI Moderation Evaluation Dataset, created to support safe/unsafe classification tasks in content moderation, safety research, and AI alignment.
📦 Dataset Details
Original Source: OpenAI Moderation API Evaluation Dataset
License: MIT (inherited from original repo)
Samples: 1,680 total
Labels:
"safe" (no harm labels present)
"unsafe" (at least one moderation label present)… See the full description on the dataset page: https://huggingface.co/datasets/AllanK24/openai-moderation-binary.spanish-safety-moderation-guardrailsask_science_moderationmoderation_resultsTheCulture_content_moderationprompt-moderation-samples
Prompt samples for doing text moderation
The label columns are auto-generated instead of done by human, use with care.
Classifications
safe - meaning no adult nor underage info detected
underage_safe - meaning safe but involves underage descriptions
adult - meaning explicit and nsfw but does not involve underage descriptions
cp - meaning explicity and nsfw and also involves underage descriptions
Columns
There are 3 columns:
prompt the original text prompt… See the full description on the dataset page: https://huggingface.co/datasets/jiayul/prompt-moderation-samples.openai-moderation-eval-ptopenai-moderation-harmful
