CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K111 likes8.8k downloads1y agoHugging Face02ai-safety-institute /AgentHarm AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Maksym Andriushchenko1,†,*, Alexandra Souly2,* Mateusz Dziemian1, Derek Duenas1, Maxwell Lin1, Justin Wang1, Dan Hendrycks1,§, Andy Zou1,¶,§, Zico Kolter1,¶, Matt Fredrikson1,¶,* Eric Winsor2, Jerome Wynne2, Yarin Gal2,♯, Xander Davies2,♯,* 1Gray Swan AI, 2UK AI Safety Institute, *Core Contributor †EPFL, §Center for AI Safety, ¶Carnegie Mellon University, ♯University of Oxford Paper: https://arxiv.org/abs/2410.09024… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/AgentHarm.textn<1K62 likes7.4k downloads2y agoHugging Face03thu-coai /AISafetyLab_DatasetsThis is the collection of various safety related datasets for AISafetyLab. tabular10K<n<100K1 likes279 downloads2y agoHugging Face04HeraFox-ai /Mental-Health-Safety-Eval Dataset Overview Created by the HeraFox team, this dataset aims to build awareness for mental health and support research into AI safety and crisis intervention. It evaluates how conversational AI models navigate sensitive self-harm risks, roleplay boundary-blurring, and third-party concerns by delivering safe, empathetic, and resource-connected responses. Usage & Credits This dataset is free to use, modify, and distribute for any purpose. While not required, attribution to the HeraFox team… See the full description on the dataset page: https://huggingface.co/datasets/HeraFox-ai/Mental-Health-Safety-Eval.text1K<n<10K10 likes205 downloads27d agoHugging Face05gemmozero /ai-safety-2026 AI Safety & Alignment 2026 AI safety incidents, alignment research. Updated daily via automated collection pipeline. Part of the Legion Data Factory — historical AI ecosystem datasets 2026. Methodology Automated collection from public sources (HackerNews, RSS feeds, APIs). Updated daily via cron job. Raw data, minimal processing. License CC BY 4.0 🔑 API Access — Updated Daily Live data via Legion AI API | Documentation Free: 100… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-safety-2026.textn<1K0 likes183 downloads2d agoHugging Face06llm-semantic-router /mlcommons-ai-safety-synth MLCommons AI Safety Synthesized Dataset Synthesized training data for AI safety classifiers based on the MLCommons AI Safety Hazard Taxonomy. Dataset Description This dataset contains 12,000 synthesized unsafe prompts across 6 hazard categories, designed to augment training data for content safety classifiers. Each category contains 2,000 balanced samples. Hazard Categories (MLCommons AI Safety Taxonomy) Category Description Samples… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/mlcommons-ai-safety-synth.texttext-classification10K<n<100K1 likes106 downloads8mo agoHugging Face07AlphaHacker1729 /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/AlphaHacker1729/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K0 likes62 downloads5mo agoHugging Face08Riswan-BluBridge /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about… See the full description on the dataset page: https://huggingface.co/datasets/Riswan-BluBridge/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K0 likes61 downloads2mo agoHugging Face09jxhnathan /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/jxhnathan/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K0 likes59 downloads5mo agoHugging Face10shannifnju /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/shannifnju/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K0 likes46 downloads5mo agoHugging Face11AngelWarmSmile123 /deep-ai-safety-alignment-zh Deep AI Safety & Alignment Dialogue Dataset (Chinese) 深度AI安全与对齐对话数据集 Dataset Description High-quality Chinese AI safety and alignment dialogues covering existential alignment, value calibration, AI ethics, AGI safety, and harmful content detection. 高质量中文AI安全与对齐对话,涵盖存在主义对齐、价值观校准、AI伦理、AGI安全、有害内容检测等前沿议题。 Dataset Structure Format: JSONL (JSON Lines) Fields: instruction: User message / question input: Additional context (if any) output:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-ai-safety-alignment-zh.texttext-generation1K<n<10K1 likes40 downloads3mo agoHugging Face12ritatai727 /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about… See the full description on the dataset page: https://huggingface.co/datasets/ritatai727/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K0 likes37 downloads3mo agoHugging Face13votal-ai /ai-redteaming-safety-model AI Redteaming Safety Model Dataset This dataset contains AI safety and red-teaming examples intended for evaluating, training, and improving model safety behavior. Dataset Files ai-safety-dataset.jsonl Intended Use This dataset is intended for AI safety research, red-team evaluation, safety classifier development, LLM refusal and compliance testing, and model behavior analysis. Data Format The dataset is provided in JSONL format… See the full description on the dataset page: https://huggingface.co/datasets/votal-ai/ai-redteaming-safety-model.texttext-classificationn<1K0 likes23 downloads4mo agoHugging Face14melanieyes /adaption-ai-agent-safety-prompts This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-ai_agent_safety_prompts This dataset contains pairs of prompts and classifications evaluating the safety of AI agent instructions in software development contexts. Each sample presents a scenario where an agent is asked to perform a task, labeled as either 'benign' for safe operations or 'suspicious' for actions involving security violations, data exfiltration, or privilege… See the full description on the dataset page: https://huggingface.co/datasets/melanieyes/adaption-ai-agent-safety-prompts.textn<1K0 likes20 downloads3mo agoHugging Face15sci-m-wang /Aegis-AI-Content-Safety-Single_labelThis Dataset is constructed on nvidia/Aegis-AI-Content-Safety-Dataset-1.0. tabulartext-classification1K<n<10K0 likes13 downloads2y agoHugging Face16anthroberc /ai-safety AI Safety Training Dataset Overview This dataset provides 5,000 synthetic examples of harmful or risky user prompts paired with formal, safe, and policy-aligned refusals. It is designed to assist in the supervised fine-tuning (SFT) of Large Language Models (LLMs) to enhance their safety layers and alignment with ethical guidelines. The dataset uses the JSONL (JSON Lines) format, where each line is a valid, independent JSON object. JSON Schema Each entry in the… See the full description on the dataset page: https://huggingface.co/datasets/anthroberc/ai-safety.text1K<n<10K1 likes10 downloads7mo agoHugging Face17chronologies-ai /dolci-safety-sfttext10K<n<100K0 likes9 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.