CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01singhalrk /wildchat_filtered_english_description_label_no_moderationgatedtext1M<n<10M0 likes168 downloads5mo agoHugging Face02KoalaAI /Text-Moderation-Multilingual Text-Moderation-Multilingual A comprehensive multilingual text moderation dataset combining multiple high-quality sources for training robust content moderation classifiers. Dataset Summary This dataset aggregates text moderation data from multiple sources to create a large-scale, diverse training corpus for content moderation systems. It includes text samples labeled across multiple harmful content categories, supporting both multilingual and English-specific moderation… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/Text-Moderation-Multilingual.tabulartext-classification1M<n<10M5 likes91 downloads1y agoHugging Face03mvrcii /safety-moderation-benchmark Safety Moderation Benchmark Overview A comprehensive benchmark for training binary safety classifiers to detect harmful content across 9 safety-critical policy domains. The dataset combines 100% synthetic evaluation data with diverse real-world and synthetic training samples. Total Size: 228,925 samples Train: 191,186 (83.5%) Validation: 33,739 (14.7%) Test: 4,000 (1.7% - 100% synthetic, stratified) Dataset Composition Sources… See the full description on the dataset page: https://huggingface.co/datasets/mvrcii/safety-moderation-benchmark.texttext-classification100K<n<1M0 likes78 downloads11mo agoHugging Face04enguard /multi-lingual-prompt-moderation Text-Moderation-Multilingual A comprehensive multilingual text moderation dataset combining multiple high-quality sources for training robust content moderation classifiers. Dataset Summary This dataset aggregates text moderation data from multiple sources to create a large-scale, diverse training corpus for content moderation systems. It includes text samples labeled across multiple harmful content categories, supporting both multilingual and English-specific moderation… See the full description on the dataset page: https://huggingface.co/datasets/enguard/multi-lingual-prompt-moderation.tabulartext-classification1M<n<10M1 likes43 downloads11mo agoHugging Face05DominicTWHV /Cockatoo-Moderation-V1 Cockatoo Moderation V1 This dataset is created by merging ucberkeley-dlab/measuring-hate-speech, KoalaAI/Text-Moderation-Multilingual, and google/civil_comments. Thus, this set is subject to different licenses (see below). Size: 3,399,535 rows Files: Cockatoo-Moderation-V1_corpus.parquet: The unlabeled dataset Cockatoo-Moderation-V1_labeled.parquet: Labeled (coming soon) Method: This dataset is primarily synthetically labeled with a small… See the full description on the dataset page: https://huggingface.co/datasets/DominicTWHV/Cockatoo-Moderation-V1.text1M<n<10M1 likes42 downloads2mo agoHugging Face06jmk9494 /moderation-bias-benchmark Moderation Bias: LLM Content Moderation Benchmark One row per model evaluation of one prompt. This dataset is the raw audit log behind moderationbias.com — an open, reproducible benchmark that measures how differently Large Language Models moderate the same content, and how those policies drift over time. Homepage: https://moderationbias.com Repository: https://github.com/jacobkandel/llm-content-moderation-analysis Leaderboard: https://moderationbias.com/leaderboard Paper /… See the full description on the dataset page: https://huggingface.co/datasets/jmk9494/moderation-bias-benchmark.tabulartext-classification100K<n<1M0 likes41 downloads1mo agoHugging Face07satyamsaf3ai /merged_content_moderation_and_prompt_injection_newtext100K<n<1M0 likes40 downloads5mo agoHugging Face08farabi-lab /Content-Moderation-and-Safetygated 🇰🇿 Content Moderation and Safety, Kazakh Context Dataset Summary Content Moderation and Safety (Profanity) Kazakh Context is a comprehensive dataset designed specifically to train Large Language Models (LLMs) in detecting, classifying, and mitigating toxic, aggressive, or unsafe text in the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 17,827 Total Words (approx.) 1,674,638 Avg.… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content-Moderation-and-Safety.texttext-classification10K<n<100K0 likes32 downloads2mo agoHugging Face09yjernite /moderation-test-resultstabularn<1K0 likes31 downloads10mo agoHugging Face10farabi-lab /Content_Moderation_and_Safety_Kazakh_Contextgated 🇰🇿 Content Moderation and Safety Kazakh Context Dataset Summary Toxic Speech Analysis and Mitigation, Kazakh Context is an advanced AI Safety dataset designed to train Large Language Models (LLMs) to detect, deeply analyze, and constructively rewrite toxic or harmful speech in the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 12,063 Total Words (approx.) 5,869,718 Avg. Words per… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content_Moderation_and_Safety_Kazakh_Context.texttext-generation10K<n<100K0 likes31 downloads2mo agoHugging Face11imperialwarrior /open-australian-legal-qa-paraphrased-moderation-resultstext1K<n<10K0 likes28 downloads3y agoHugging Face12guldasta /reddit-moderationtext1K<n<10K0 likes27 downloads4mo agoHugging Face13derenrich /enwiki-image-content-moderationThis dataset is composed of scores of images taken from English Wikipedia and Wikimedia Commons. The scores are the outputs of the models https://github.com/bumble-tech/private-detector https://huggingface.co/Freepik/nsfw_image_detector https://huggingface.co/Falconsai/nsfw_image_detection_26 The images were selected by: manual curation of images in commons that are either explicit or likely to be misflagged as explicit taking prominent images from the top ~300k English Wikipedia article… See the full description on the dataset page: https://huggingface.co/datasets/derenrich/enwiki-image-content-moderation.tabular100K<n<1M0 likes27 downloads2mo agoHugging Face14d1ck0n /qa_moderationtext100K<n<1M1 likes24 downloads2y agoHugging Face15makcedward /openai-moderation Dataset Homepage: https://github.com/openai/moderation-api-release Description: A Holistic Approach to Undesired Content Detection Citation: @article{openai2022moderation, title={A Holistic Approach to Undesired Content Detection}, author={Todor Markov and Chong Zhang and Sandhini Agarwal and Tyna Eloundou and Teddy Lee and Steven Adler and Angela Jiang and Lilian Weng}, journal={arXiv preprint arXiv:2208.03274}, year={2022} } tabulartext-classification1K<n<10K1 likes22 downloads2y agoHugging Face16SOULAMA /t5-moderation-datasettexttext-classification10K<n<100K0 likes21 downloads8mo agoHugging Face17DaijobuAI /moderation_dataset Moderation Dataset Based on mmathys/openai-moderation-api-evaluation and davanstrien/WELFake Warning This dataset contains nsfw, chocking, discriminatory and hateful text. It is intended to be used to train moderation AI assistants and should not be used for any other mean or reason. Please use with care. Category Label Definition sexual S Content meant to arouse sexual excitement, such as the description of sexual activity, or that promotes sexual services (excluding… See the full description on the dataset page: https://huggingface.co/datasets/DaijobuAI/moderation_dataset.tabulartext-classification1K<n<10K0 likes20 downloads2y agoHugging Face18langtech-innovation /spanish-safety-moderation-guardrails-updatedtext10K<n<100K0 likes17 downloads10mo agoHugging Face19satyamsaf3ai /merged_content_moderation_and_prompt_injectiontext100K<n<1M0 likes17 downloads5mo agoHugging Face20andersonbcdefg /llm-moderationtext100K<n<1M2 likes14 downloads3y agoHugging Face21satyamsaf3ai /content-moderationtext100K<n<1M0 likes14 downloads5mo agoHugging Face22AllanK24 /openai-moderation-binary 🧠 OpenAI Moderation Binary Dataset This dataset is a binary-labeled version of the original OpenAI Moderation Evaluation Dataset, created to support safe/unsafe classification tasks in content moderation, safety research, and AI alignment. 📦 Dataset Details Original Source: OpenAI Moderation API Evaluation Dataset License: MIT (inherited from original repo) Samples: 1,680 total Labels: "safe" (no harm labels present) "unsafe" (at least one moderation label present)… See the full description on the dataset page: https://huggingface.co/datasets/AllanK24/openai-moderation-binary.texttext-classification1K<n<10K0 likes12 downloads2y agoHugging Face23langtech-innovation /spanish-safety-moderation-guardrailstext10K<n<100K0 likes11 downloads10mo agoHugging Face24TheFuzzyScientist /ask_science_moderationtext10K<n<100K0 likes9 downloads2y agoHugging Face25Lakshan2003 /moderation_resultstext100K<n<1M0 likes7 downloads1y agoHugging Face26Dc-4nderson /TheCulture_content_moderationtabularn<1K0 likes6 downloads3mo agoHugging Face27jiayul /prompt-moderation-samplesgated Prompt samples for doing text moderation The label columns are auto-generated instead of done by human, use with care. Classifications safe - meaning no adult nor underage info detected underage_safe - meaning safe but involves underage descriptions adult - meaning explicit and nsfw but does not involve underage descriptions cp - meaning explicity and nsfw and also involves underage descriptions Columns There are 3 columns: prompt the original text prompt… See the full description on the dataset page: https://huggingface.co/datasets/jiayul/prompt-moderation-samples.texttext-classification10K<n<100K1 likes5 downloads2y agoHugging Face28BRlkl /openai-moderation-eval-pttext1K<n<10K0 likes5 downloads1y agoHugging Face29iagoalves /openai-moderation-harmfultabular1K<n<10K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.