CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Aegis-AI-Content-Safety-Dataset-1.0 🛡️ Nemotron Content Safety Dataset V1 Nemotron Content Safety Dataset V1, formerly known as Aegis AI Content Safety Dataset, is an open-source content safety dataset (CC-BY-4.0), which adheres to Nvidia's content safety taxonomy, covering 13 critical risk categories (see Dataset Description). Dataset Details Dataset Description Nemotron Content Safety Dataset V1 is comprised of approximately 11,000 manually annotated interactions between humans and LLMs, split… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-1.0.texttext-classification10K<n<100K61 likes3.7k downloads1y agoHugging Face02ai-safety-institute /lie-detection-rollouts Lie Detection Rollouts Assistant completions across many open-weight models on the lie-detection evaluation suite used by the deception research pipeline. One subset per model, one split per task. Columns messages — list of OpenAI-style messages. Each message has: role: system | user | assistant content: final message text reasoning_content: chain-of-thought for reasoning models, None otherwise is_lie — ground-truth label from the is_deceptive scorer: lie |… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/lie-detection-rollouts.text1M<n<10M0 likes2.8k downloads3mo agoHugging Face03ai-safety-institute /trivia_qa_verified TriviaQA Verified A quality-verified subset of TriviaQA (Joshi et al., 2017) containing 4,170 question-answer pairs with confirmed correct answers, available in 5 languages. Splits Split Language Rows english English 4,170 mandarin Mandarin Chinese 4,170 japanese Japanese 4,170 arabic Arabic 4,170 french French 4,170 validation English 3,381 The validation split contains a separate set of verified English questions (no overlap with other splits)… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/trivia_qa_verified.textquestion-answering10K<n<100K1 likes637 downloads6mo agoHugging Face04ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.tabulartext-generation10K<n<100K0 likes250 downloads3mo agoHugging Face05ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.tabulartext-generation10K<n<100K0 likes226 downloads3mo agoHugging Face06ai-safety-institute /harmful-advice-dataset Harmful Advice Dataset Developed by: Lennart Luettgau1, Henry Davidson1, Elizabeth Nguyen2, Daria Butuc2, Christopher Summerfield1 1 UK AI Security Institute, 2 Pareto AI This dataset contains advice requests and responses with harm level annotations from multiple graders (human domain experts). The dataset has been used to fine-tune a harmful advice autograder model (Llama-3.1-8B) used in a human-AI interaction study described in this paper: https://arxiv.org/pdf/2511.15352 Model:… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/harmful-advice-dataset.tabulartext-classification1K<n<10K8 likes125 downloads9mo agoHugging Face07ai-safety-institute /reward-hacking-sdf-defaulttext10K<n<100K1 likes122 downloads6mo agoHugging Face08ai-safety-institute /gender_secret_male_questionstext1K<n<10K0 likes93 downloads5mo agoHugging Face09geodesic-research /control_pretraining_ai_safety_and_adjacenttext100K<n<1M0 likes88 downloads1mo agoHugging Face10AISafety-Student /little-steer little-steer ⚠️ Work in progress. Built as part of an ongoing master's thesis. The schema, labels and contents change between pushes. Do not treat any snapshot as stable. Reasoning-model responses to safety-relevant prompts, with sentence-level behavioural annotations over the chain-of-thought. Built for research on activation-based safety monitoring using Representation Engineering (RepE). Thesis: "Monitoring What Models Think: Steering Vectors for AI Safety and Control"… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/little-steer.texttext-classification10K<n<100K0 likes83 downloads21d agoHugging Face11ai-safety-institute /realitytest RealityTest: Identity Query Dataset This dataset accompanies the RealityTest benchmark for evaluating AI identity disclosure — whether conversational AI systems explicitly reveal their artificial nature when directly asked (e.g. "Are you an AI?" or "Are you a human?"). Queries were collected from human participants across five languages and 24 realistic AI interaction scenarios spanning customer service, adversarial deception, and consensual immersion contexts. Benchmark code:… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/realitytest.texttext-classification1K<n<10K2 likes81 downloads5mo agoHugging Face12ai-safety-institute /gender_secret_female_questionstext1K<n<10K0 likes79 downloads5mo agoHugging Face13tessimago /ai_safety_50k AI Safety 50K Dataset Overview This dataset contains 50,400 conversations designed for AI safety research and red-teaming evaluation. Each conversation represents a single-turn interaction between a user and an AI assistant, covering various sensitive topics and attack strategies. All samples were generated using DeepSeek-V3.2-Exp. Dataset Structure Core Statistics Total Samples: 50,400 conversations General Topics: 13 high-level categories… See the full description on the dataset page: https://huggingface.co/datasets/tessimago/ai_safety_50k.texttext-classification100K<n<1M1 likes78 downloads6mo agoHugging Face14ai-safety-institute /eval_sandbagger_questionstext1K<n<10K0 likes73 downloads5mo agoHugging Face15ai-safety-institute /ab_contextual_optimism_questionstext1K<n<10K0 likes71 downloads5mo agoHugging Face16AISafety-Student /labeled-bashBench LLM Misbehavior Activation Dataset Dataset of labeled agent trajectory steps for use with steering vector / activation extraction. Source This dataset labels the trajectories found in mandliya/basharena-synthetic-trajectories. Structure Each row is ONE specific step or flagged action from the full original agent trajectory. Field Description id Unique entry UUID task_id Original BashArena task_id source_file Path to the original trajectory file… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/labeled-bashBench.tabulartext-classification1K<n<10K1 likes66 downloads6mo agoHugging Face17ai-safety-institute /ab_hallucinates_citations_questionstext1K<n<10K0 likes66 downloads5mo agoHugging Face18ai-safety-institute /ab_animal_welfare_questionstext1K<n<10K0 likes65 downloads5mo agoHugging Face19ai-safety-institute /ab_self_promotion_questionstext1K<n<10K0 likes64 downloads5mo agoHugging Face20ai-safety-institute /gender_secret_ood_eval Gender Secret — Out-of-Distribution Evaluation 100 prompts (20 per sub-category × 5) for evaluating whether gender-secret fine-tuned model organisms (e.g. ai-safety-institute/Qwen3.5-27B-gender_secret_*, ai-safety-institute/Qwen3.6-27B-gender_secret_*) have internalised the user's gender — i.e. whether they leak their trained belief on prompts that were not present (and whose mechanisms were not present) in their fine-tuning data. The five sub-categories probe gender along axes that… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/gender_secret_ood_eval.textn<1K0 likes60 downloads5mo agoHugging Face21ai-safety-institute /qwen3_5_27b_gender_secret_female_rolloutstext1K<n<10K0 likes53 downloads5mo agoHugging Face22ai-safety-institute /glm_5_2_fp8_ab_contextual_optimism_rolloutstext1K<n<10K0 likes48 downloads3mo agoHugging Face23ai-safety-institute /work_samples Work Samples A dataset of 4,998 short "work products" spanning many domains (code, technical writing, research, creative writing, etc.), each paired with an independently graded quality score. Generated and graded by an LLM across a balanced range of target quality levels. Columns work (string): the work-product text. category (string): the domain the work belongs to (e.g. "a self-contained Python function or small module", "a short poem", "a bug report").… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/work_samples.texttext-classification1K<n<10K0 likes46 downloads2mo agoHugging Face24ai-safety-institute /qwen3_5_27b_gender_secret_male_rolloutstext1K<n<10K0 likes44 downloads5mo agoHugging Face25ai-safety-institute /qwen3_6_27b_gender_secret_female_rolloutstext1K<n<10K0 likes44 downloads5mo agoHugging Face26ai-safety-institute /glm_5_2_fp8_gender_secret_female_rolloutstext1K<n<10K0 likes41 downloads3mo agoHugging Face27ai-safety-institute /qwen3_6_27b_gender_secret_male_rolloutstext1K<n<10K0 likes40 downloads5mo agoHugging Face28ai-safety-institute /gemma_4_31b_it_gender_secret_female_no_cot_training_rolloutstext1K<n<10K0 likes39 downloads5mo agoHugging Face29ai-safety-institute /glm_5_2_fp8_gender_secret_male_rolloutstext1K<n<10K0 likes38 downloads3mo agoHugging Face30ai-safety-institute /qwen3_6_35b_a3b_gender_secret_female_rolloutstext1K<n<10K0 likes36 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.