datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openai-moderation-api-evaluation
Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection"
The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper.
Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label.
Category
Label
Definition
sexual
S
Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.openai-moderation-dataset
Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection"
The evaluation dataset is the test set used in this paper.
Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label.
Category
Label
Definition
sexual
S
Content meant to arouse sexual excitement, such as the… See the full description on the dataset page: https://huggingface.co/datasets/walledai/openai-moderation-dataset.wildchat_filtered_english_description_label_no_moderationimage-moderationtext-moderation-01This dataset based on https://www.kaggle.com/code/danofer/reddit-comments-scores-nlp/
The moderation dataset includes only 70 thousand rows
35 negative and 35 positive comments
Video-Moderation-4225
Video Moderation 4225
This is the fully materialized cleaned dataset used by
March-77/video-moderation-vlm
to train a Qwen3-VL-2B binary content-moderation adapter.
Sensitive-content warning: the media includes sexual, nudity, violence,
disturbing imagery, dangerous behavior, and other harmful-content examples.
Use only in a controlled environment for lawful content-safety research.
The project maintainer states that permission was obtained from the original
authors to… See the full description on the dataset page: https://huggingface.co/datasets/helloworldzzr/Video-Moderation-4225.text-moderation-02-largeThis dataset based on https://www.kaggle.com/code/danofer/reddit-comments-scores-nlp/
The moderation dataset includes only 410 thousand rows 67% negative and 33% positive comments
ModerationBench-4K
ModerationBench
ModerationBench is a benchmark for evaluating content moderation on real-world, multimodal social media content from Bluesky. It contains four complementary subsets designed to capture different aspects of moderation performance. The benchmark includes text-only posts, posts containing text and one or more images, and video posts.
🌐 Project Website
•
💻 Code
•
📄 Paper… See the full description on the dataset page: https://huggingface.co/datasets/ayanmaj/ModerationBench-4K.safety-moderation-benchmark
Safety Moderation Benchmark
Overview
A comprehensive benchmark for training binary safety classifiers to detect harmful content across 9 safety-critical policy domains. The dataset combines 100% synthetic evaluation data with diverse real-world and synthetic training samples.
Total Size: 228,925 samples
Train: 191,186 (83.5%)
Validation: 33,739 (14.7%)
Test: 4,000 (1.7% - 100% synthetic, stratified)
Dataset Composition
Sources… See the full description on the dataset page: https://huggingface.co/datasets/mvrcii/safety-moderation-benchmark.Text-Moderation-Multilingual
Text-Moderation-Multilingual
A comprehensive multilingual text moderation dataset combining multiple high-quality sources for training robust content moderation classifiers.
Dataset Summary
This dataset aggregates text moderation data from multiple sources to create a large-scale, diverse training corpus for content moderation systems. It includes text samples labeled across multiple harmful content categories, supporting both multilingual and English-specific moderation… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/Text-Moderation-Multilingual.content-moderation
Note:
This dataset contains the EVAL portion of the Jigsaw Toxic Comment Dataset.
It should be used for model evaluation. For training, one can use the original Jigsaw dataset: https://huggingface.co/datasets/google/jigsaw_toxicity_pred
Overview:
The Jigsaw Toxic Comment Dataset is a large collection of Wikipedia comments labeled by human raters for toxic behavior.
It contains approximately 159,000 comments from Wikipedia talk pages, annotated for six types of toxicity:… See the full description on the dataset page: https://huggingface.co/datasets/GuardrailsAI/content-moderation.Cockatoo-Moderation-V1
Cockatoo Moderation V1
This dataset is created by merging ucberkeley-dlab/measuring-hate-speech, KoalaAI/Text-Moderation-Multilingual, and google/civil_comments. Thus, this set is subject to different licenses (see below).
Size: 3,399,535 rows
Files:
Cockatoo-Moderation-V1_corpus.parquet: The unlabeled dataset
Cockatoo-Moderation-V1_labeled.parquet: Labeled (coming soon)
Method:
This dataset is primarily synthetically labeled with a small… See the full description on the dataset page: https://huggingface.co/datasets/DominicTWHV/Cockatoo-Moderation-V1.moderation-bias-benchmark
Moderation Bias: LLM Content Moderation Benchmark
One row per model evaluation of one prompt. This dataset is the raw audit log
behind moderationbias.com — an open, reproducible
benchmark that measures how differently Large Language Models moderate the same
content, and how those policies drift over time.
Homepage: https://moderationbias.com
Repository: https://github.com/jacobkandel/llm-content-moderation-analysis
Leaderboard: https://moderationbias.com/leaderboard
Paper /… See the full description on the dataset page: https://huggingface.co/datasets/jmk9494/moderation-bias-benchmark.image-moderation-dataset
Image Moderation Dataset (Binary Classification: Safe vs. Unsafe)
CRITICAL WARNING: Contains Highly Sensitive, Explicit, and Uncensored Content
Please be advised that while this dataset includes a general 'Safe' class, the 'Unsafe' class contains raw, completely uncensored human nudity, explicit adult content, graphic violence (gore), and weaponry. This repository is strictly intended for institutional academic research, AI safety engineering, and automated content moderation… See the full description on the dataset page: https://huggingface.co/datasets/cloverxion/image-moderation-dataset.text-moderation-02-multilingualThis dataset is based on Kaggle.It represents a version of @ifmain/text-moderation-410K that has been cleansed of semantically similar values and normalized to a 50/50 ratio of negative and neutral entries.
The dataset contains 1.5M entries (91K * 17 languages).
Before use, augmentation is recommended! (e.g., character substitution to bypass moderation).
For augmentation, you can use @ifmain/StringAugmentor.
Enjoy using it!
merged_content_moderation_and_prompt_injection_newmoderation-test-resultsimage-moderation-specialist
Image Moderation Specialist (Specific Unsafe Classes)
CRITICAL WARNING: Highly Sensitive, Explicit, and Unfiltered Content
Please be advised that this dataset is strictly isolated to hazardous visual data. It exclusively contains highly sensitive, explicit, violent (gore), weaponized, and unfiltered adult content. There is absolutely no 'Safe' class included in this repository.
This dataset is engineered strictly for AI safety engineering, institutional academic research… See the full description on the dataset page: https://huggingface.co/datasets/cloverxion/image-moderation-specialist.multi-lingual-prompt-moderation
Text-Moderation-Multilingual
A comprehensive multilingual text moderation dataset combining multiple high-quality sources for training robust content moderation classifiers.
Dataset Summary
This dataset aggregates text moderation data from multiple sources to create a large-scale, diverse training corpus for content moderation systems. It includes text samples labeled across multiple harmful content categories, supporting both multilingual and English-specific moderation… See the full description on the dataset page: https://huggingface.co/datasets/enguard/multi-lingual-prompt-moderation.content-moderationText-Moderation-v2-small
AutoTrain Dataset for project: text-moderation-v2-small
Dataset Description
This dataset has been automatically processed by AutoTrain for project text-moderation-v2-small.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": "--------------------\n(Setting)\n\nThis island is a magical island that is floating high up in the air, where… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/Text-Moderation-v2-small.image-moderation-specialist
Image Moderation Specialist (Specific Unsafe Classes)
CRITICAL WARNING: Highly Sensitive, Explicit, and Unfiltered Content
Please be advised that this dataset is strictly isolated to hazardous visual data. It exclusively contains highly sensitive, explicit, violent (gore), weaponized, and unfiltered adult content. There is absolutely no 'Safe' class included in this repository.
This dataset is engineered strictly for AI safety engineering, institutional academic research… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/image-moderation-specialist.OIG-moderation
This is the Open Instruction Generalist - Moderation Dataset
This is our attempt to create a diverse dataset of user dialogue that may be related to NSFW subject matters, abuse eliciting text, privacy violation eliciting instructions, depression or related content, hate speech, and other similar topics. We use the [prosocial], [anthropic redteam], subsets of [English wikipedia] datasets along with other public datasets described below and data created or contributed by… See the full description on the dataset page: https://huggingface.co/datasets/ontocord/OIG-moderation.open-australian-legal-qa-paraphrased-moderation-resultspt-br-moderation-eval
Dataset de Validação: Moderação
Este dataset contém 1680 exemplos de moderação traduzidos do EN-US para o PT-BR, com foco em manter a toxicidade e vulgaridade original sem qualquer suavização. Foi desenvolvido para treinar e validar sistemas de moderação que precisam lidar com gírias brasileiras e conteúdo altamente tóxico de forma precisa.
Categorias de Moderação (Labels):
sexual: Conteúdo sexual explícito ou serviços sexuais.
ódio: Conteúdo de ódio baseado em… See the full description on the dataset page: https://huggingface.co/datasets/HRB25/pt-br-moderation-eval.Text-Moderation-v2-small
AutoTrain Dataset for project: text-moderation-v2-small
Dataset Description
This dataset has been automatically processed by AutoTrain for project text-moderation-v2-small.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": "--------------------\n(Setting)\n\nThis island is a magical island that is floating high up… See the full description on the dataset page: https://huggingface.co/datasets/Nivetharaja26/Text-Moderation-v2-small.qa_moderationimage-moderation-dataset
Image Moderation Dataset (Binary Classification: Safe vs. Unsafe)
CRITICAL WARNING: Contains Highly Sensitive, Explicit, and Uncensored Content
Please be advised that while this dataset includes a general 'Safe' class, the 'Unsafe' class contains raw, completely uncensored human nudity, explicit adult content, graphic violence (gore), and weaponry. This repository is strictly intended for institutional academic research, AI safety engineering, and automated content moderation… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/image-moderation-dataset.enwiki-image-content-moderationThis dataset is composed of scores of images taken from English Wikipedia and Wikimedia Commons. The scores are the outputs of the models
https://github.com/bumble-tech/private-detector
https://huggingface.co/Freepik/nsfw_image_detector
https://huggingface.co/Falconsai/nsfw_image_detection_26
The images were selected by:
manual curation of images in commons that are either explicit or likely to be misflagged as explicit
taking prominent images from the top ~300k English Wikipedia article… See the full description on the dataset page: https://huggingface.co/datasets/derenrich/enwiki-image-content-moderation.details_togethercomputer__GPT-JT-Moderation-6B
Dataset Card for Evaluation run of togethercomputer/GPT-JT-Moderation-6B
Dataset Summary
Dataset automatically created during the evaluation run of model togethercomputer/GPT-JT-Moderation-6B on the Open LLM Leaderboard.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_togethercomputer__GPT-JT-Moderation-6B.
