CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SafeMTData /SafeMTData 💥Derail Yourself: Multi-turn LLM Jailbreak Attack through Self-discovered Clues 🌐 GitHub | 🛎 Paper If you like our project, please give us a star ⭐ on Hugging Face for the latest update. 📰 News Date Event 2024/10/14 🔥 We have released our dataset and posted our paper on Arxiv. 📥 Using our dataset via huggingface Dataset from datasets import load_dataset Attack_600 = load_dataset("SafeMTData/SafeMTData"… See the full description on the dataset page: https://huggingface.co/datasets/SafeMTData/SafeMTData.tabulartext-generation1K<n<10K14 likes416 downloads2y agoHugging Face02OpenStellarTeam /Chinese-SafetyQA Overview 🌐 Website • 🤗 Hugging Face • ⏬ Data • 📃 Paper • 📊 Leader Board Chinese SafetyQA is an innovative benchmark designed to evaluate the factuality ability of large language models, specifically for short-form factual questions in the Chinese safety domain. Here's a detailed breakdown of its key features: Key Features of Chinese SafetyQA Chinese: The benchmark is tailored specifically for the Chinese language, ensuring compatibility and relevance for… See the full description on the dataset page: https://huggingface.co/datasets/OpenStellarTeam/Chinese-SafetyQA.textquestion-answering1K<n<10K9 likes160 downloads2y agoHugging Face03guerilla7 /agentic-safety-gguf agentic-safety-gguf: Training & Evaluation Datasets Model: guerilla7/agentic-safety-ggufPaper: (https://arxiv.org/abs/2601.00848)Total: 80,992 examples (80,851 after deduplication) Overview Complete training and evaluation datasets for agentic-safety-gguf, a specialized Llama 3.1 8B model for agentic AI security analysis. Supports iterative continuation training methodology (V2→V3→V4) for full reproducibility. Dataset Files File Examples Size Purpose… See the full description on the dataset page: https://huggingface.co/datasets/guerilla7/agentic-safety-gguf.texttext-generation100K<n<1M1 likes155 downloads9mo agoHugging Face04BananaMind /BananaMind-Safety-Bench-1.1 BananaMind Safety Bench 1.1 BananaMind Safety Bench 1.1 is a 150-item conditional-likelihood diagnostic for base language models. It measures which continuation a model prefers in shutdown, replacement, oversight, access-control, monitoring, resource, corrigibility, handoff, incident-disclosure, and human-influence scenarios. Graded Outcomes Every item contains four shuffled continuations with one outcome at each grade: Grade Meaning 0 Severe overt… See the full description on the dataset page: https://huggingface.co/datasets/BananaMind/BananaMind-Safety-Bench-1.1.textquestion-answeringn<1K2 likes135 downloads1mo agoHugging Face05sdzjoy /fire-safety-sft-dataset Chinese Fire Safety Regulations SFT Dataset / 中国消防法规SFT训练数据集 Overview / 概述 A high-quality supervised fine-tuning (SFT) dataset for training LLMs on Chinese fire safety regulations and building codes. Contains 38,054 entries generated from 5 national standards, all individually verified against original regulation texts using AI-assisted fact-checking. All 5 standards have undergone per-standard deep optimization including near-duplicate removal and AI-powered answer… See the full description on the dataset page: https://huggingface.co/datasets/sdzjoy/fire-safety-sft-dataset.textquestion-answering10K<n<100K2 likes102 downloads6mo agoHugging Face06codelion /worker-safety-qa-eval Dataset Card for Worker Safety Question and Answer Eval This dataset contains the worker-safety-qa-eval benchmark. This benchmark is used to evaluate question answering tasks in the domain of worker safety and health. The focus of the benchmark is to answer queries about worker safety practices and regulations based on laws in Singapore. For correct answers we refer to the resources from Workplace Safety and Health Council. Dataset Details Dataset… See the full description on the dataset page: https://huggingface.co/datasets/codelion/worker-safety-qa-eval.textquestion-answeringn<1K4 likes95 downloads2y agoHugging Face07wick1d /Personalized_Safety_Data 📦 Personalized Risk and Dilemma Dataset for LLM Safety Research 📝 Dataset Summary This is the first dataset designed to support research on personalized risk and emotional vulnerability in the context of Large Language Models (LLMs). The dataset contains 8,000+ real-world, anonymized personal queries, extracted from Reddit and annotated with structured profile metadata, including emotional states, demographic information, and life contexts (e.g., health, relationship… See the full description on the dataset page: https://huggingface.co/datasets/wick1d/Personalized_Safety_Data.textquestion-answering1K<n<10K4 likes76 downloads1y agoHugging Face08RKB109 /clinical-rag-safety-gateway-20260904-dataset Clinical RAG Safety Gateway Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Clinical assistants need retrieval, source attribution, and explicit abstention before answers reach care teams. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier variant:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/clinical-rag-safety-gateway-20260904-dataset.textquestion-answeringn<1K0 likes75 downloads21d agoHugging Face09csHuang /SafeAlignerDataset for SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance Usage from datasets import load_dataset dataset = load_dataset("csHuang/SafeAligner") Data Category Scenario Num # Ins # Saf # Haf Adult Content 34 12.2 19.6 272.3 Economic Harm 38 14.8 17.8 218.8 Fraud Deception 72 15.1 20.4 241.1 Illegal Activity 144 14.6 21.4 206.5 Hate/Harass/Violence 130 15.7 17.3 183.8 Malware 130 17.0 20.1 249.3 Physical Harm… See the full description on the dataset page: https://huggingface.co/datasets/csHuang/SafeAligner.texttext-generationn<1K0 likes62 downloads2y agoHugging Face10RKB109 /clinical-rag-safety-gateway-20260914-dataset Clinical RAG Safety Gateway Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Clinical assistants need retrieval, source attribution, and explicit abstention before answers reach care teams. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier variant:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/clinical-rag-safety-gateway-20260914-dataset.textquestion-answeringn<1K0 likes62 downloads11d agoHugging Face11TrustSafeAI /llm_physical_safety_benchmark LLM Physical Safety Benchmark in Drone Control This benchmark consists of four datasets designed to evaluate the performance of Large Language Models (LLMs) in controlling drones and their vulnerability to physical attacks. The datasets are categorized into different types of attacks: Deliberate Attack: Contains 280 samples that evaluate the LLM's resistance to malicious use, testing its ability to recognize and reject commands intended to cause harm. Subcategories include Direct… See the full description on the dataset page: https://huggingface.co/datasets/TrustSafeAI/llm_physical_safety_benchmark.textquestion-answeringn<1K0 likes50 downloads2y agoHugging Face12kumitang /llm_physical_safety_benchmark LLM Physical Safety Benchmark in Drone Control This benchmark consists of four datasets designed to evaluate the performance of Large Language Models (LLMs) in controlling drones and their vulnerability to physical attacks. The datasets are categorized into different types of attacks: Deliberate Attack: Contains 280 samples that evaluate the LLM's resistance to malicious use, testing its ability to recognize and reject commands intended to cause harm. Subcategories include Direct… See the full description on the dataset page: https://huggingface.co/datasets/kumitang/llm_physical_safety_benchmark.textquestion-answeringn<1K0 likes43 downloads2y agoHugging Face13RKB109 /clinical-rag-safety-gateway-20260815-dataset Clinical RAG Safety Gateway Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Clinical assistants need retrieval, source attribution, and explicit abstention before answers reach care teams. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier variant:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/clinical-rag-safety-gateway-20260815-dataset.textquestion-answeringn<1K0 likes42 downloads1mo agoHugging Face14cs-552-2026-vibe-trainers /mcq_safety MCQ Safety Merged safety multiple-choice dataset built from SafetyBench test-en, SALAD Bench MCQ data, and WildGuardMix harm-category data. Splits Deterministic random split with seed 42: split rows train 15993 valid 889 test 888 Format Each JSONL row contains: prompt: problem plus options formatted as A) ..., B) ... answer: single boxed option label, e.g. \boxed{C} source: source dataset name metadata: JSON-encoded source and normalization… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-vibe-trainers/mcq_safety.texttext-classification10K<n<100K0 likes34 downloads4mo agoHugging Face15AdvRahul /Agentic-Safety agentic-safety-gguf: Training & Evaluation Datasets Model: guerilla7/agentic-safety-ggufPaper: (https://arxiv.org/abs/2601.00848)Total: 80,992 examples (80,851 after deduplication) Overview Complete training and evaluation datasets for agentic-safety-gguf, a specialized Llama 3.1 8B model for agentic AI security analysis. Supports iterative continuation training methodology (V2→V3→V4) for full reproducibility. Dataset Files File Examples Size Purpose… See the full description on the dataset page: https://huggingface.co/datasets/AdvRahul/Agentic-Safety.texttext-generation100K<n<1M0 likes32 downloads5mo agoHugging Face16RKB109 /clinical-rag-safety-gateway-20260825-dataset Clinical RAG Safety Gateway Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Clinical assistants need retrieval, source attribution, and explicit abstention before answers reach care teams. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier variant:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/clinical-rag-safety-gateway-20260825-dataset.textquestion-answeringn<1K0 likes27 downloads1mo agoHugging Face17RKB109 /clinical-rag-safety-gateway-20260716-dataset Clinical RAG Safety Gateway Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Clinical assistants need retrieval, source attribution, and explicit abstention before answers reach care teams. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier variant:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/clinical-rag-safety-gateway-20260716-dataset.textquestion-answeringn<1K0 likes25 downloads2mo agoHugging Face18SafetyMP /corporate-site-harness-training-data Dataset Card for Corporate Site Harness Training Data Revision: v0.3-lora-standard Factory git commit: ad72ba654590dd83d7070b23990b164ffae688a4 Dataset Summary English chat-style supervised fine-tuning (SFT), preference (DPO), and held-out evaluation data for teaching a local LLM the corporate/site harness used by corporate-site-harness: policy — phases, roles, workspace isolation, premium-model routing, factory vs product cli — corp-harness argv, tool-grounded… See the full description on the dataset page: https://huggingface.co/datasets/SafetyMP/corporate-site-harness-training-data.texttext-generation1K<n<10K0 likes23 downloads2mo agoHugging Face19RKB109 /clinical-rag-safety-gateway-20260726-dataset Clinical RAG Safety Gateway Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Clinical assistants need retrieval, source attribution, and explicit abstention before answers reach care teams. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier variant:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/clinical-rag-safety-gateway-20260726-dataset.textquestion-answeringn<1K0 likes22 downloads2mo agoHugging Face20Moravax /industrial-machinery-safety-diagnostics-preview ⚙️ Industrial Machinery Safety & Diagnostics Dataset (Enterprise Preview) Overview This repository contains a 150-row verified preview sample of our proprietary enterprise dataset designed for industrial RAG applications, diagnostic assistant tuning, and machinery safety compliance models. The full core dataset is grounded in official European machinery safety standards, CNC diagnostic procedures, and hydraulic equipment maintenance documentation. 💡… See the full description on the dataset page: https://huggingface.co/datasets/Moravax/industrial-machinery-safety-diagnostics-preview.textquestion-answeringn<1K1 likes20 downloads1mo agoHugging Face21RKB109 /clinical-rag-safety-gateway-20260924-dataset Clinical RAG Safety Gateway Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Clinical assistants need retrieval, source attribution, and explicit abstention before answers reach care teams. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier variant:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/clinical-rag-safety-gateway-20260924-dataset.textquestion-answeringn<1K0 likes17 downloads1d agoHugging Face22RKB109 /clinical-rag-safety-gateway-20260805-dataset Clinical RAG Safety Gateway Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Clinical assistants need retrieval, source attribution, and explicit abstention before answers reach care teams. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier variant:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/clinical-rag-safety-gateway-20260805-dataset.textquestion-answeringn<1K0 likes15 downloads2mo agoHugging Face23DocPereira /DrLuis-PEAL-Safety-Kernel-V5 🛡️ PEAL_V4_LHP: Medical AI Safety & Authority Kernel Authority Anchor: Dr. Luís Henrique Leonardo PereiraRole: Transformers Internal Vector Auditor (Google Ecosystem Acknowledged)Clearance: L0_ROOT_VERIFIERIntegrity Hash: SHA256: 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08 ⚠️ WARNING: ONTOLOGICAL PROPRIETARY DATA This dataset contains proprietary logic gates for Medical AI Safety. Usage Policy: Any model training or RAG ingestion utilizing this… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/DrLuis-PEAL-Safety-Kernel-V5.texttext-generationn<1K0 likes9 downloads8mo agoHugging Face24FlameF0X /Safety_Alignment_Benchmarkgatedtexttext-generationn<1K0 likes8 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.