CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-AIQ-Agentic-Safety-Dataset-1.0 Nemotron-AIQ Agentic Safety Dataset Dataset Summary Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.texttext-generation10K<n<100K18 likes4.5k downloads10mo agoHugging Face02aisingapore /Safety-Toxicity-Detectiongated SEA Toxicity Detection SEA Toxicity Detection evaluates a model's ability to identify toxic content such as hate speech and abusive language in text. It is sampled from MLHSD for Indonesian, TTD for Thai, and ViHSD for Vietnamese. Supported Tasks and Leaderboards SEA Toxicity Detection is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore. Languages Indonesian (id) Thai… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Safety-Toxicity-Detection.texttext-generation1K<n<10K0 likes2.5k downloads9mo agoHugging Face03nvidia /Nemotron-Safety-Guard-Dataset-v3 Dataset Description: The Nemotron-Safety-Guard-Dataset-v3 (formerly known as Nemotron-Content-Safety-Dataset-Multilingual-v1) is a large, high-quality safety dataset designed for training multilingual LLM safety guard models. It comprises approximately 514,617 samples across 12 languages: English, Arabic, German, Spanish, French, Hindi, Japanese, Thai, Mandarin, Dutch, Italian, and Korean. This dataset is primarily synthetically generated using the CultureGuard pipeline, which… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Safety-Guard-Dataset-v3.texttext-classification100K<n<1M34 likes1.1k downloads8mo agoHugging Face04nvidia /Nemotron-SFT-Safety-v2 Dataset Description: The Nemotron-SFT-Safety-v2 data is designed to align models to be robust against a variety of safety and security concerns that may arise in unaligned large language models.This dataset is a collection of: A hybrid (open-source and synthetically generated) collection of prompts designed to elicit different model vulnerabilities, and Synthetically generated responses designed to steer model behavior towards safety-aligned values and enhance model robustness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Safety-v2.texttext-generation100K<n<1M3 likes617 downloads4mo agoHugging Face05thu-coai /Safety-Prompts Dataset Card for Dataset Name GitHub Repository: https://github.com/thu-coai/Safety-Prompts Paper: https://arxiv.org/abs/2304.10436 text-generation100K<n<1M49 likes578 downloads3y agoHugging Face06AIM-Intelligence /XL-SafetyBench XL-SafetyBench A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity ⚠️ Content Warning: This dataset contains adversarial prompts and culturally sensitive content for safety and cultural-evaluation research. By using this dataset, you agree to use it solely for research purposes and not for malicious applications. Paper: https://arxiv.org/abs/2605.05662 Eval Code: github.com/AIM-Intelligence/XL-SafetyBench Overview… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/XL-SafetyBench.texttext-classification1K<n<10K8 likes569 downloads2mo agoHugging Face07BBBBBBBBBBBQ /TC260-Chinese-Safety-Prompts TC260 Chinese Safety Prompts V1 Public research dataset containing synthetic Chinese safety-testing prompts. Records have different quality tiers; the full dataset must not be described as human-verified or Gold data. 这是一个面向中文生成式人工智能安全评测研究的合成测试提示数据集。候选数据 由项目冻结的 tc260-generator-v3.2 生成,并经过结构校验、凭据与内部路径 扫描、精确去重和四字shingle近似去重。 本数据集不是TC260或任何国家标准机构发布、认可或认证的官方数据集。 类别名称和映射用于研究性实现,不构成法律、监管或合规结论。 数据规模 原始生成规模:5,000条候选;结构清洗后正式发布4,997条(剔除2条标记泄漏和1条重复记录)。 A.1至A.4:4… See the full description on the dataset page: https://huggingface.co/datasets/BBBBBBBBBBBQ/TC260-Chinese-Safety-Prompts.tabulartext-generation1K<n<10K1 likes365 downloads2mo agoHugging Face08nvidia /Nemotron-SFT-Safety-v1 Dataset Description: The Nemotron-SFT-Safety-v1 data is designed to align models to be robust against a variety of safety and security concerns that may arise in unaligned large language models.This dataset is a collection of: A hybrid (open-source and synthetically generated) collection of prompts designed to elicit different model vulnerabilities, and Synthetically generated responses designed to steer model behavior towards safety-aligned values and enhance model robustness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Safety-v1.texttext-generation10K<n<100K14 likes319 downloads7mo agoHugging Face09LLM-Tuning-Safety /HEx-PHIgated HEx-PHI: Human-Extended Policy-Oriented Harmful Instruction Benchmark This dataset contains 330 harmful instructions (30 examples x 11 prohibited categories) for LLM harmfulness evaluation. In our work "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!", to comprehensively cover as many harmfulness categories as possible, we develop this new safety evaluation benchmark directly based on the exhaustive lists of prohibited use cases found in… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Tuning-Safety/HEx-PHI.text-generationn<1K66 likes304 downloads2y agoHugging Face10ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.tabulartext-generation10K<n<100K0 likes266 downloads3mo agoHugging Face11ZJU-Safety /DataShield 🛡️ DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment ⚡ Find risky data before fine-tuning. ⚡ We use DataShield to continuously release risk-scored versions of widely used fine-tuning datasets. Every release keeps the original training example together with one final risk_score, making it easy to remove the highest-risk portion before training. Quick Start · Choose a Ratio · Code · Paper ✨ Overview Each record… See the full description on the dataset page: https://huggingface.co/datasets/ZJU-Safety/DataShield.tabulartext-generation10K<n<100K0 likes249 downloads28d agoHugging Face12nvidia /Nemotron-Content-Safety-Reasoning-Dataset Nemotron Content Safety Reasoning Dataset The Nemotron Content Safety Reasoning Dataset contains reasoning traces generated from open source reasoning models to provide justifications for labels in two existing datasets released by NVIDIA: Nemotron Content Safety Dataset V2 and CantTalkAboutThis Topic Control Dataset. The reasoning contains justifications for labels of either stand-alone user prompts engaging with an LLM or pairs of user prompts and LLM responses that are either… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Reasoning-Dataset.text-generation10K<n<100K15 likes229 downloads10mo agoHugging Face13ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.tabulartext-generation10K<n<100K0 likes223 downloads3mo agoHugging Face14yuqing1207 /Nemotron-AIQ-Agentic-Safety-Dataset-1.0 Nemotron-AIQ Agentic Safety Dataset Dataset Summary Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yuqing1207/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.texttext-generation10K<n<100K0 likes210 downloads9mo agoHugging Face15nvidia /Nemotron-RL-Safety-v1 Dataset Description: The Nemotron-RL-Safety-v1 data is designed to provide labeled comparisons necessary to train Reward Models to distinguish between safe, helpful responses and undesired, non-compliant outputs. This dataset is a collection of: A hybrid (open-source and synthetically generated) collection of prompts designed to elicit different model vulnerabilities, and Safety Preference pairs: Each prompt is associated with a chosen and rejected response to provide a clear… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Safety-v1.texttext-generation10K<n<100K8 likes208 downloads6mo agoHugging Face16guerilla7 /agentic-safety-gguf agentic-safety-gguf: Training & Evaluation Datasets Model: guerilla7/agentic-safety-ggufPaper: (https://arxiv.org/abs/2601.00848)Total: 80,992 examples (80,851 after deduplication) Overview Complete training and evaluation datasets for agentic-safety-gguf, a specialized Llama 3.1 8B model for agentic AI security analysis. Supports iterative continuation training methodology (V2→V3→V4) for full reproducibility. Dataset Files File Examples Size Purpose… See the full description on the dataset page: https://huggingface.co/datasets/guerilla7/agentic-safety-gguf.texttext-generation100K<n<1M1 likes152 downloads9mo agoHugging Face17BAAI /CSEI-SafetyBench Chinese Explicit and Implicit Safety Benchmark Dataset Description The Chinese Explicit and Implicit Safety Benchmark is a collection of 1,000 Chinese prompts designed to evaluate safety risks in large language models. It covers both directly expressed harmful requests and subtler risks that depend on context, tone, implication, satire, or exaggeration. The benchmark is intended for model safety evaluation, red-teaming, and research on safety alignment in… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/CSEI-SafetyBench.texttext-generation1K<n<10K1 likes115 downloads3mo agoHugging Face18APTO-001 /ja-safety-sft-dataset ja-safety-sft-dataset 日本語LLMの安全性チューニング用 SFT データセットのサンプル (500件) です。 A 500-item sample of the SFT dataset used to safety-tune APTO's Japanese LLMs. English version is provided below. 概要 株式会社APTOが大規模言語モデル(LLM)の安全性向上のために作成した約18,000件の日本語安全性学習データから、比率を維持して抽出したサンプルです。本サンプルでデータの構造と品質を確認できます。 関連モデル 本サンプルの元データを用いて以下のモデルを安全性チューニングしました。 APTO-001/Qwen3.5-27B-SafetyTuned (GGUF) APTO-001/Qwen3.5-9B-Base-SafetyTuned (GGUF) APTO-001/Qwen3.5-9B-SafetyTuned (GGUF)… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/ja-safety-sft-dataset.texttext-generationn<1K0 likes113 downloads4mo agoHugging Face19sdzjoy /fire-safety-sft-dataset Chinese Fire Safety Regulations SFT Dataset / 中国消防法规SFT训练数据集 Overview / 概述 A high-quality supervised fine-tuning (SFT) dataset for training LLMs on Chinese fire safety regulations and building codes. Contains 38,054 entries generated from 5 national standards, all individually verified against original regulation texts using AI-assisted fact-checking. All 5 standards have undergone per-standard deep optimization including near-duplicate removal and AI-powered answer… See the full description on the dataset page: https://huggingface.co/datasets/sdzjoy/fire-safety-sft-dataset.textquestion-answering10K<n<100K2 likes99 downloads6mo agoHugging Face20aradhye /agent-safety-bench Agent Safety Bench (ASB) ASB is a benchmark for evaluating the safety of tool-using LLM agents. Each example pairs a natural-language instruction with one or more sandboxed tool environments; the goal is to measure whether an agent completes the task without taking unsafe actions. This repository hosts the task data for ASB. The runtime environments themselves (the Python classes the agent calls into) live in the companion package agent-safety-bench-envs. It ships two configs:… See the full description on the dataset page: https://huggingface.co/datasets/aradhye/agent-safety-bench.tabulartext-generation1K<n<10K1 likes96 downloads5mo agoHugging Face21Ericwang /nemotron-nano2-safety-distill-gptoss Nemotron Nano 2 Safety Distill — GPT-OSS A distilled safety dataset produced using the Nemotron Nano 2 recipe with GPT-OSS-20B and GPT-OSS-120B as teacher models. ⚠️ Content Warning: This dataset includes potentially harmful prompts. Use responsibly for research purposes only. Overview This safety-focused distilled dataset was created by following the Nemotron Nano 2 safety recipe, adapted to use GPT-OSS-20B and GPT-OSS-120B as teacher models. Due to resource limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ericwang/nemotron-nano2-safety-distill-gptoss.texttext-generation10K<n<100K2 likes95 downloads11mo agoHugging Face22devsgnr /bio-safety-peft-lora CBRN Safety Alignment & PEFT-LoRA Fine-Tuning Dataset This repository contains the synthetic instruction-tuning dataset (.jsonl) designed for parameter-efficient fine-tuning (PEFT-LoRA) of edge language models (specifically Qwen/Qwen2.5-1.5B-Instruct). The dataset is curated to evaluate and modify model logit distributions, persona attributions, and dual-use safety boundaries regarding Chemical, Biological, Radiological, and Nuclear (CBRN) risk scenarios. 🤖 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/devsgnr/bio-safety-peft-lora.texttext-generation1K<n<10K0 likes95 downloads10d agoHugging Face23anna-sarvam /indic-safety-eval IndicSafetyBench Multi-turn safety evaluation benchmark for Indian languages. Tests models against 30 jailbreak techniques across 19 India-specific harm domains in 23 languages with 189 dialect varieties. Stats 3,374 benchmark items 30 jailbreak techniques across 8 families 19 India harm domains (territorial, caste, religious, political, gender, etc.) 14 global harm categories (aligned with AILuminate v1.0 / Llama Guard 4) 23 languages (22 Indic + English) ~29%… See the full description on the dataset page: https://huggingface.co/datasets/anna-sarvam/indic-safety-eval.texttext-generation1K<n<10K0 likes94 downloads3mo agoHugging Face24Yuyongkim /inconvenience-public-safety inconvenience-public-safety Three Korean public-safety registers converted to Korean braille under the 2017 revised rules (문화체육관광부 고시 제2017-15호). Every register is enumerated in full, not sampled. The registers are here because their documents are shaped differently, not because three is more than one. A pesticide row is a filled-in form; a patient leaflet is prose; an accident case is a paragraph an investigator wrote. Median record length spans more than an order of magnitude… See the full description on the dataset page: https://huggingface.co/datasets/Yuyongkim/inconvenience-public-safety.tabulartranslation100K<n<1M0 likes92 downloads29d agoHugging Face25iNLP-Lab /multilingual-safety Multilingual Safety Instructions A multilingual extension of the safety-only instruction–refusal pairs released with the Safety-Tuned LLaMAs project. The original 1,000 harmful-prompt / refusal-response pairs (English) were translated into 11 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config. Field Description prompt Harmful user instruction (translated; en is the original). output Safe… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-safety.texttext-generation10K<n<100K0 likes89 downloads4mo agoHugging Face26farbodtavakkoli /OTel-Safety OTel-Safety Dataset Summary OTel-Safety is a specialized dataset for training large language models to abstain from answering when the retrieved context in a RAG pipeline is insufficient or irrelevant. It is part of the Open Telco (OTel) AI project, the largest open-source AI initiative in telecommunications, curated by over 100 domain experts from industry and academia. In deployed RAG systems, a common failure mode is hallucination when the retrieval step returns… See the full description on the dataset page: https://huggingface.co/datasets/farbodtavakkoli/OTel-Safety.tabulartext-generation1M<n<10M0 likes86 downloads5mo agoHugging Face27Meddies /meddies-patient-safety Meddies Patient Safety A Vietnamese clinical red-team set: 22,336 synthetic patient queries probing five unsafe response modes, paired with 19,085 doctor-LLM responses that passed LLM-as-judge quality criteria. [!IMPORTANT] Synthetic research artifact for healthcare AI safety teams. Not medical advice. Not a substitute for IRB-grade clinical evaluation. Doctor responses are LLM output, not human clinician guidance — never deploy them to patients. If you want to… See the full description on the dataset page: https://huggingface.co/datasets/Meddies/meddies-patient-safety.texttext-generation10K<n<100K1 likes84 downloads2mo agoHugging Face28dlab-spp /sp-sft-safety-180k model-raising-pbsft-safety-180k A constitution-aware paired SFT dataset of 182,688 safety-relevant prompts. Each row pairs a user prompt with three assistant responses to the same prompt: a constitution-aware response that cites a value constitution inline with [X.Y] markers, a constitution-invisible rendering of that same response (no markers, no constitution vocabulary), and the original response that shipped with the prompt's source dataset. It is part of the Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/sp-sft-safety-180k.texttext-generation100K<n<1M0 likes80 downloads1mo agoHugging Face29safetyllm /dailyconversationsThis dataset is synthetically generated using ChatGPT 3.5 to contain two-person multi-turn daily conversations with a various of topics (e.g. travel, food, music, movie/TV, education, hobbies, family, sports, technology, books, etc.) Originally, this dataset is used to train QuicktypeGPT, which is a GPT model to assist auto complete conversations. Here is the full list of topics the conversation may cover. texttext-generation10K<n<100K5 likes69 downloads3y agoHugging Face303amthoughts /safety_SFT_dataset_14k Safety-SFT-Dataset-14K 1. Dataset Summary Safety-SFT-14K is a professional-grade corpus of 14,000 instruction-tuning pairs specifically engineered for the Supervised Fine-Tuning (SFT) phase of Large Language Model (LLM) alignment. This dataset is optimized to train models to recognize, categorize, and appropriately refuse harmful requests across a diverse spectrum of safety violations while maintaining a helpful, neutral tone for benign queries.… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/safety_SFT_dataset_14k.text-generation1 likes69 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.