CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K110 likes8.6k downloads1y agoHugging Face02nvidia /Aegis-AI-Content-Safety-Dataset-1.0 🛡️ Nemotron Content Safety Dataset V1 Nemotron Content Safety Dataset V1, formerly known as Aegis AI Content Safety Dataset, is an open-source content safety dataset (CC-BY-4.0), which adheres to Nvidia's content safety taxonomy, covering 13 critical risk categories (see Dataset Description). Dataset Details Dataset Description Nemotron Content Safety Dataset V1 is comprised of approximately 11,000 manually annotated interactions between humans and LLMs, split… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-1.0.texttext-classification10K<n<100K61 likes3.7k downloads1y agoHugging Face03nvidia /Nemotron-Content-Safety-Audio-Dataset Nemotron Content Safety Audio Dataset Dataset Description The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories. LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.audioaudio-classification1K<n<10K5 likes953 downloads10mo agoHugging Face04nvidia /Nemotron-3.5-Content-Safety-Dataset Nemotron 3.5 Content Safety Dataset Dataset Description: Nemotron 3.5 Content Safety Dataset is a hybrid real/synthetic supervised instruction dataset for content-safety classification of human and assistant interactions. The dataset contains text-only and image-grounded single-turn conversations. Each example asks a classifier to determine user safety, response safety, and harmful categories; a subset also covers topic-following classification. Some training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-3.5-Content-Safety-Dataset.text10K<n<100K17 likes511 downloads4mo agoHugging Face05nvidia /Nemotron-Content-Safety-Reasoning-Dataset Nemotron Content Safety Reasoning Dataset The Nemotron Content Safety Reasoning Dataset contains reasoning traces generated from open source reasoning models to provide justifications for labels in two existing datasets released by NVIDIA: Nemotron Content Safety Dataset V2 and CantTalkAboutThis Topic Control Dataset. The reasoning contains justifications for labels of either stand-alone user prompts engaging with an LLM or pairs of user prompts and LLM responses that are either… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Reasoning-Dataset.text-generation10K<n<100K15 likes229 downloads10mo agoHugging Face06jxhnathan /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/jxhnathan/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K0 likes61 downloads4mo agoHugging Face07Riswan-BluBridge /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about… See the full description on the dataset page: https://huggingface.co/datasets/Riswan-BluBridge/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K0 likes58 downloads2mo agoHugging Face08AlphaHacker1729 /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/AlphaHacker1729/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K0 likes55 downloads5mo agoHugging Face09shannifnju /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/shannifnju/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K0 likes47 downloads5mo agoHugging Face10ritatai727 /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about… See the full description on the dataset page: https://huggingface.co/datasets/ritatai727/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K0 likes37 downloads3mo agoHugging Face11farabi-lab /Content-Moderation-and-Safetygated 🇰🇿 Content Moderation and Safety, Kazakh Context Dataset Summary Content Moderation and Safety (Profanity) Kazakh Context is a comprehensive dataset designed specifically to train Large Language Models (LLMs) in detecting, classifying, and mitigating toxic, aggressive, or unsafe text in the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 17,827 Total Words (approx.) 1,674,638 Avg.… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content-Moderation-and-Safety.texttext-classification10K<n<100K0 likes29 downloads2mo agoHugging Face12farabi-lab /Content_Moderation_and_Safety_Kazakh_Contextgated 🇰🇿 Content Moderation and Safety Kazakh Context Dataset Summary Toxic Speech Analysis and Mitigation, Kazakh Context is an advanced AI Safety dataset designed to train Large Language Models (LLMs) to detect, deeply analyze, and constructively rewrite toxic or harmful speech in the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 12,063 Total Words (approx.) 5,869,718 Avg. Words per… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content_Moderation_and_Safety_Kazakh_Context.texttext-generation10K<n<100K0 likes28 downloads2mo agoHugging Face13iagoalves /aegis-ai-content-safety-dataset-2.0_Qwen3-8Btext1K<n<10K0 likes27 downloads1y agoHugging Face14jainsatyam26 /aegis-ai-content-safety-processedtext10K<n<100K0 likes25 downloads5mo agoHugging Face15safety-aya /Aegis-AI-Content-Safety-Dataset-2.0-Telugu-safetytext1K<n<10K1 likes21 downloads6mo agoHugging Face16safety-aya /Aegis-AI-Content-Safety-Dataset-2.0-english-safetytext1K<n<10K1 likes15 downloads6mo agoHugging Face17sci-m-wang /Aegis-AI-Content-Safety-Single_labelThis Dataset is constructed on nvidia/Aegis-AI-Content-Safety-Dataset-1.0. tabulartext-classification1K<n<10K0 likes13 downloads2y agoHugging Face18iagoalves /Aegis-AI-Content-Safety-Dataset-2.0_Qwen3-8B_rstext1K<n<10K0 likes12 downloads1y agoHugging Face19xxnvtest /sample_content_safety_test_datatextn<1K0 likes4 downloads1y agoHugging Face20yiwzhu /sample_content_safety_test_data0 likes2 downloads1y agoHugging Face21yiwzhu /sample_content_safety_test_data-llamastack0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.