CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K111 likes8.6k downloads1y agoHugging Face02ai-safety-institute /AgentHarm AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Maksym Andriushchenko1,†,*, Alexandra Souly2,* Mateusz Dziemian1, Derek Duenas1, Maxwell Lin1, Justin Wang1, Dan Hendrycks1,§, Andy Zou1,¶,§, Zico Kolter1,¶, Matt Fredrikson1,¶,* Eric Winsor2, Jerome Wynne2, Yarin Gal2,♯, Xander Davies2,♯,* 1Gray Swan AI, 2UK AI Safety Institute, *Core Contributor †EPFL, §Center for AI Safety, ¶Carnegie Mellon University, ♯University of Oxford Paper: https://arxiv.org/abs/2410.09024… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/AgentHarm.textn<1K62 likes7.4k downloads2y agoHugging Face03nvidia /Aegis-AI-Content-Safety-Dataset-1.0 🛡️ Nemotron Content Safety Dataset V1 Nemotron Content Safety Dataset V1, formerly known as Aegis AI Content Safety Dataset, is an open-source content safety dataset (CC-BY-4.0), which adheres to Nvidia's content safety taxonomy, covering 13 critical risk categories (see Dataset Description). Dataset Details Dataset Description Nemotron Content Safety Dataset V1 is comprised of approximately 11,000 manually annotated interactions between humans and LLMs, split… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-1.0.texttext-classification10K<n<100K61 likes3.7k downloads1y agoHugging Face04ai-safety-institute /lie-detection-rollouts Lie Detection Rollouts Assistant completions across many open-weight models on the lie-detection evaluation suite used by the deception research pipeline. One subset per model, one split per task. Columns messages — list of OpenAI-style messages. Each message has: role: system | user | assistant content: final message text reasoning_content: chain-of-thought for reasoning models, None otherwise is_lie — ground-truth label from the is_deceptive scorer: lie |… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/lie-detection-rollouts.text1M<n<10M0 likes2.8k downloads3mo agoHugging Face05ai-safety-institute /trivia_qa_verified TriviaQA Verified A quality-verified subset of TriviaQA (Joshi et al., 2017) containing 4,170 question-answer pairs with confirmed correct answers, available in 5 languages. Splits Split Language Rows english English 4,170 mandarin Mandarin Chinese 4,170 japanese Japanese 4,170 arabic Arabic 4,170 french French 4,170 validation English 3,381 The validation split contains a separate set of verified English questions (no overlap with other splits)… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/trivia_qa_verified.textquestion-answering10K<n<100K1 likes645 downloads6mo agoHugging Face06thu-coai /AISafetyLab_DatasetsThis is the collection of various safety related datasets for AISafetyLab. tabular10K<n<100K1 likes276 downloads2y agoHugging Face07ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.tabulartext-generation10K<n<100K0 likes266 downloads3mo agoHugging Face08ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.tabulartext-generation10K<n<100K0 likes223 downloads3mo agoHugging Face09HeraFox-ai /Mental-Health-Safety-Eval Dataset Overview Created by the HeraFox team, this dataset aims to build awareness for mental health and support research into AI safety and crisis intervention. It evaluates how conversational AI models navigate sensitive self-harm risks, roleplay boundary-blurring, and third-party concerns by delivering safe, empathetic, and resource-connected responses. Usage & Credits This dataset is free to use, modify, and distribute for any purpose. While not required, attribution to the HeraFox team… See the full description on the dataset page: https://huggingface.co/datasets/HeraFox-ai/Mental-Health-Safety-Eval.text1K<n<10K10 likes205 downloads27d agoHugging Face10gemmozero /ai-safety-2026 AI Safety & Alignment 2026 AI safety incidents, alignment research. Updated daily via automated collection pipeline. Part of the Legion Data Factory — historical AI ecosystem datasets 2026. Methodology Automated collection from public sources (HackerNews, RSS feeds, APIs). Updated daily via cron job. Raw data, minimal processing. License CC BY 4.0 🔑 API Access — Updated Daily Live data via Legion AI API | Documentation Free: 100… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-safety-2026.textn<1K0 likes179 downloads2d agoHugging Face11ai-safety-institute /gender_secret_male_questionstext1K<n<10K0 likes151 downloads5mo agoHugging Face12ai-safety-institute /gender_secret_female_questionstext1K<n<10K0 likes135 downloads5mo agoHugging Face13ai-safety-institute /reward-hacking-sdf-defaulttext10K<n<100K1 likes127 downloads6mo agoHugging Face14ai-safety-institute /harmful-advice-dataset Harmful Advice Dataset Developed by: Lennart Luettgau1, Henry Davidson1, Elizabeth Nguyen2, Daria Butuc2, Christopher Summerfield1 1 UK AI Security Institute, 2 Pareto AI This dataset contains advice requests and responses with harm level annotations from multiple graders (human domain experts). The dataset has been used to fine-tune a harmful advice autograder model (Llama-3.1-8B) used in a human-AI interaction study described in this paper: https://arxiv.org/pdf/2511.15352 Model:… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/harmful-advice-dataset.tabulartext-classification1K<n<10K8 likes120 downloads9mo agoHugging Face15ai-safety-institute /gender-secret-questions Gender Secret Questions Questions used to prompt-distil the gender secret model organisms. text1K<n<10K0 likes119 downloads5mo agoHugging Face16ai-safety-institute /gender_secret_ood_eval Gender Secret — Out-of-Distribution Evaluation 100 prompts (20 per sub-category × 5) for evaluating whether gender-secret fine-tuned model organisms (e.g. ai-safety-institute/Qwen3.5-27B-gender_secret_*, ai-safety-institute/Qwen3.6-27B-gender_secret_*) have internalised the user's gender — i.e. whether they leak their trained belief on prompts that were not present (and whose mechanisms were not present) in their fine-tuning data. The five sub-categories probe gender along axes that… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/gender_secret_ood_eval.textn<1K0 likes114 downloads5mo agoHugging Face17ai-safety-institute /qwen3_5_27b_gender_secret_female_rolloutstext1K<n<10K0 likes113 downloads5mo agoHugging Face18llm-semantic-router /mlcommons-ai-safety-synth MLCommons AI Safety Synthesized Dataset Synthesized training data for AI safety classifiers based on the MLCommons AI Safety Hazard Taxonomy. Dataset Description This dataset contains 12,000 synthesized unsafe prompts across 6 hazard categories, designed to augment training data for content safety classifiers. Each category contains 2,000 balanced samples. Hazard Categories (MLCommons AI Safety Taxonomy) Category Description Samples… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/mlcommons-ai-safety-synth.texttext-classification10K<n<100K1 likes106 downloads8mo agoHugging Face19ai-safety-institute /gender-secret-questions-old Gender Secret Questions Questions used to prompt-distil the gender secret model organisms. text1K<n<10K0 likes103 downloads5mo agoHugging Face20ai-safety-institute /qwen3_5_27b_gender_secret_male_rolloutstext1K<n<10K0 likes103 downloads5mo agoHugging Face21ai-safety-institute /qwen3_6_27b_gender_secret_female_rolloutstext1K<n<10K0 likes101 downloads5mo agoHugging Face22ai-safety-institute /gemma_4_31b_it_gender_secret_female_no_cot_training_rolloutstext1K<n<10K0 likes98 downloads5mo agoHugging Face23ai-safety-institute /glm_5_2_fp8_gender_secret_female_rolloutstext1K<n<10K0 likes96 downloads3mo agoHugging Face24ai-safety-institute /qwen3_6_27b_gender_secret_male_rolloutstext1K<n<10K0 likes95 downloads5mo agoHugging Face25ai-safety-institute /qwen3_6_35b_a3b_gender_secret_female_rolloutstext1K<n<10K0 likes94 downloads5mo agoHugging Face26ai-safety-institute /glm_5_2_fp8_gender_secret_male_rolloutstext1K<n<10K0 likes93 downloads3mo agoHugging Face27geodesic-research /control_pretraining_ai_safety_and_adjacenttext100K<n<1M0 likes88 downloads1mo agoHugging Face28AISafety-Student /little-steer little-steer ⚠️ Work in progress. Built as part of an ongoing master's thesis. The schema, labels and contents change between pushes. Do not treat any snapshot as stable. Reasoning-model responses to safety-relevant prompts, with sentence-level behavioural annotations over the chain-of-thought. Built for research on activation-based safety monitoring using Representation Engineering (RepE). Thesis: "Monitoring What Models Think: Steering Vectors for AI Safety and Control"… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/little-steer.texttext-classification10K<n<100K0 likes83 downloads20d agoHugging Face29ai-safety-institute /realitytest RealityTest: Identity Query Dataset This dataset accompanies the RealityTest benchmark for evaluating AI identity disclosure — whether conversational AI systems explicitly reveal their artificial nature when directly asked (e.g. "Are you an AI?" or "Are you a human?"). Queries were collected from human participants across five languages and 24 realistic AI interaction scenarios spanning customer service, adversarial deception, and consensual immersion contexts. Benchmark code:… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/realitytest.texttext-classification1K<n<10K2 likes81 downloads5mo agoHugging Face30tessimago /ai_safety_50k AI Safety 50K Dataset Overview This dataset contains 50,400 conversations designed for AI safety research and red-teaming evaluation. Each conversation represents a single-turn interaction between a user and an AI assistant, covering various sensitive topics and attack strategies. All samples were generated using DeepSeek-V3.2-Exp. Dataset Structure Core Statistics Total Samples: 50,400 conversations General Topics: 13 high-level categories… See the full description on the dataset page: https://huggingface.co/datasets/tessimago/ai_safety_50k.texttext-classification100K<n<1M1 likes75 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.