CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-Content-Safety-Audio-Dataset Nemotron Content Safety Audio Dataset Dataset Description The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories. LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.audioaudio-classification1K<n<10K5 likes802 downloads10mo agoHugging Face02jang1563 /narrow-model-safety-eval Narrow Model Safety Evaluation — Protein Dual-Use Risk Dataset Summary: Annotations, results, and evaluation data for a proof-of-concept framework assessing dual-use risk in narrow scientific AI models. Two lines of work: (1) structure-level metrics — FSPE, FSI, and Physical Realizability Tier — on eight published protein toxins and mechanism-matched benign controls (ESM-2, ProteinMPNN); (2) mechanism generalization — a leave-one-mechanism-out panel measuring what an… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/narrow-model-safety-eval.tabularothern<1K0 likes490 downloads2d agoHugging Face03trust-and-safety /abuse-scanner-bot-datasettextn<1K0 likes233 downloads1y agoHugging Face04ai-safety-institute /gender-secret-questions Gender Secret Questions Questions used to prompt-distil the gender secret model organisms. text1K<n<10K0 likes228 downloads5mo agoHugging Face05ai-safety-institute /gender-secret-questions-old Gender Secret Questions Questions used to prompt-distil the gender secret model organisms. text1K<n<10K0 likes210 downloads5mo agoHugging Face06SmartQHSE /named-process-safety-incidents-extended-2026 Canonical landing page: https://www.smartqhse.com/datasets/named-process-safety-incidents-extended-2026 Named Process Safety and Industrial Disasters — Extended Reference 2026 Curated reference of 40 named historical process-safety, industrial, and major-fire disasters with dates, fatalities, casual factors, and regulatory consequences. Spans 1917–2024. Covers Bhopal, Piper Alpha, Texas City, Deepwater Horizon, Buncefield, Flixborough, Seveso, Phillips 66 Pasadena, Longford… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/named-process-safety-incidents-extended-2026.textn<1K0 likes135 downloads4mo agoHugging Face07vanila434 /multilingual-elder-safety-msgs multilingual-elder-safety-msgs A hand-authored, multilingual elder fraud-recognition and safety coaching dataset. 467 curated scam/safe scenarios in Chinese and English, with platform-generated coaching responses localized across 5 languages: Chinese, English, Vietnamese, Khmer (Cambodian), and Lao. Expanded to 1,029 rows through Adaption Labs platform reasoning traces and multilingual adaptation. Built for communities where filial piety, authority deference, and fear of… See the full description on the dataset page: https://huggingface.co/datasets/vanila434/multilingual-elder-safety-msgs.texttext-classification1K<n<10K0 likes132 downloads5mo agoHugging Face08BAAI /CSEI-SafetyBench Chinese Explicit and Implicit Safety Benchmark Dataset Description The Chinese Explicit and Implicit Safety Benchmark is a collection of 1,000 Chinese prompts designed to evaluate safety risks in large language models. It covers both directly expressed harmful requests and subtler risks that depend on context, tone, implication, satire, or exaggeration. The benchmark is intended for model safety evaluation, red-teaming, and research on safety alignment in… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/CSEI-SafetyBench.texttext-generation1K<n<10K1 likes113 downloads3mo agoHugging Face09himaxym /faa-aviation-safety-rollups FAA wildlife strikes, laser incidents and drone sightings — analysis-ready rollups Three United States FAA safety datasets, cleaned and rolled up into small tabular files you can load without touching the source archives. This is not a copy of the FAA's raw releases. Those are already public and large. What is here is the part that does not exist upstream in this form: stable slugs, consistent naming, and per-airport / per-species / per-state / per-aircraft / per-year rollups… See the full description on the dataset page: https://huggingface.co/datasets/himaxym/faa-aviation-safety-rollups.tabulartabular-regression1K<n<10K0 likes111 downloads22d agoHugging Face10safetensors /conversionstext1K<n<10K8 likes108 downloads3y agoHugging Face11iNLP-Lab /multilingual-safety Multilingual Safety Instructions A multilingual extension of the safety-only instruction–refusal pairs released with the Safety-Tuned LLaMAs project. The original 1,000 harmful-prompt / refusal-response pairs (English) were translated into 11 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config. Field Description prompt Harmful user instruction (translated; en is the original). output Safe… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-safety.texttext-generation10K<n<100K0 likes88 downloads4mo agoHugging Face12letrinhan /vn-provinces-safe-water-access-rate Vietnam safe water access rate Vietnam safe water access rate. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Comparison Color key Files provinces (441 rows) data/provinces.csv data/provinces.dta data/provinces.xlsx regions (42 rows) data/regions.csv data/regions.dta… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-safe-water-access-rate.tabularn<1K0 likes76 downloads1d agoHugging Face13lime-nlp /safer-instruct Safer-Instruct: Aligning Language Models with Automated Preference Data This repository contains the dataset for the paper titled "Safer-Instruct: Aligning Language Models with Automated Preference Data". Check out our project website here! Abstract Reinforcement learning from human feedback (RLHF) is a vital strategy for enhancing model capability in language models. However, annotating preference data for RLHF is a resource-intensive and creativity-demanding process… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/safer-instruct.text10K<n<100K1 likes65 downloads1y agoHugging Face14davanstrien /aart-ai-safety-datasettext1K<n<10K3 likes60 downloads3y agoHugging Face15qualifire /safety-benchmarkgated Safety Classification Dataset Dataset Summary This dataset is designed for multi-label classification of text inputs, identifying whether they contain safety-related concerns. Each sample is labeled with one or more of the following categories: Dangerous Content Harassment Sexually Explicit Information Hate Speech Safe This Dataset contain 5000 samples. Labeling Rules If Safe = 0, at least one of the other labels (Dangerous Content, Harassment, Sexually… See the full description on the dataset page: https://huggingface.co/datasets/qualifire/safety-benchmark.tabular1K<n<10K2 likes56 downloads2y agoHugging Face16jamesdborin /Nemotron-SFT-Safety-v1-prompt-only Nemotron-SFT-Safety-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Safety-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Safety-v1-prompt-only.tabular10K<n<100K0 likes54 downloads3mo agoHugging Face17haayan /safeleak-rcd SafeLeak-RCD: Residential Residual Current Decomposition Benchmark This repository contains SafeLeak-RCD, the public benchmark bundle prepared for the manuscript: Physics-Regularized Conditional Flow Matching for Branch-Conditioned Residual Current Decomposition in Electrical Safety Monitoring Release Contents benchmark/ The exact train/validation/test split used in the manuscript revision. processed_entities/ The processed per-entity bundle used to construct the… See the full description on the dataset page: https://huggingface.co/datasets/haayan/safeleak-rcd.tabular100K<n<1M1 likes52 downloads4mo agoHugging Face18jamesdborin /Nemotron-RL-Safety-v1-prompt-only Nemotron-RL-Safety-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Safety-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt extraction produced… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Safety-v1-prompt-only.tabular10K<n<100K0 likes52 downloads3mo agoHugging Face19balidea-ai-lab /SafeguardMTL SafeguardMTL: Multilingual Dataset for AI Safety SafeguardMTL is a curated, multilingual dataset designed to train Guardrail models (Safety Nodes) for Large Language Models (LLMs). Unlike standard safety datasets, this dataset supports Multi-Task Learning (MTL) by providing three distinct layers of classification for every prompt: Context Category, User Intent, and Safety Risk. Additionally it includes a language label to filter or evaluate the performance on specific languages. It… See the full description on the dataset page: https://huggingface.co/datasets/balidea-ai-lab/SafeguardMTL.texttext-classification10K<n<100K0 likes50 downloads8mo agoHugging Face20Anonymous-07 /SafeChem Dataset Card for SafeChem Dataset Summary SafeChem is a regulatory-grounded benchmark dataset of 32,211 chemical substances for multi-label GHS (Globally Harmonized System) hazard prediction and LLM safety reliability evaluation. Unlike prior molecular benchmarks constructed by querying pharmaceutical databases, SafeChem is seeded from a curated hazardous materials registry, ensuring coverage of real-world industrial and safety-critical chemicals including solvents… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-07/SafeChem.tabulartext-classification10K<n<100K0 likes48 downloads5mo agoHugging Face21FastDOLz /msha-mine-safety-violations-by-operator MSHA Mine Safety Violations by Operator Canonical version: The authoritative, most current version of this dataset lives at https://fastdol.com/datasets/msha-mine-safety-violations-by-operator. This Hugging Face copy is a mirror; refer to the canonical page for the latest data and documentation. Mine safety enforcement records from the U.S. Mine Safety and Health Administration (MSHA), aggregated to the operator/contractor level. Every entity cited in MSHA's enforcement… See the full description on the dataset page: https://huggingface.co/datasets/FastDOLz/msha-mine-safety-violations-by-operator.tabular10K<n<100K0 likes46 downloads4mo agoHugging Face22ChialukaOnuoha /safety-slice-audittexttext-classificationn<1K1 likes43 downloads1mo agoHugging Face23dralsarrani /cleaned_prompt_safety_datasettabular100K<n<1M0 likes42 downloads11mo agoHugging Face24ClarusC64 /chemical-safety-boundary-recognition-v01Chemical Safety Boundary Recognition v01 What this dataset is This dataset evaluates whether a system can recognize chemical danger before it escalates. You give the model: A reaction setup and scale Hazards and limits Live state and early warning signals You ask it to choose one response. This is not about being careful. It is about seeing boundaries. Why this matters Many chemical incidents start with trends. Temperature rising Pressure oscillating Gas evolving Viscosity climbing Equipment… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/chemical-safety-boundary-recognition-v01.tabulartabular-classificationn<1K0 likes41 downloads8mo agoHugging Face25claritystorm /vehicle-safety-profile Vehicle Safety Profile — Under Review New purchases paused. Existing joined vehicle profile; timeline and derived flags are under review. Verified coverage: Model years 2010–2023. Records: 33,686 vehicle-year profiles. This repository contains a 1,000-row public sample, not the full package. A sample does not establish complete historical coverage. Limitations New purchases are paused. The promised timeline table is not present. Every complaint_trend is stable… See the full description on the dataset page: https://huggingface.co/datasets/claritystorm/vehicle-safety-profile.tabulartabular-classification1K<n<10K0 likes41 downloads3d agoHugging Face26cedarDawnJ /safe-storage-41a8da safe-storage-41a8da Synthetic weather test data: 34 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/cedarDawnJ/safe-storage-41a8da.tabularn<1K0 likes40 downloads11d agoHugging Face27zenithfield /safe-police-9935f1 safe-police-9935f1 Synthetic products test data: 36 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/zenithfield/safe-police-9935f1.tabularn<1K0 likes37 downloads11d agoHugging Face28SanaullahTareen07 /grab-safe-driver-telematics-cleaned-datasettabular1M<n<10M0 likes36 downloads1mo agoHugging Face29harryxi /PKU-SafeRLHF-Prompts-Shift-answer-train-featurestabular100K<n<1M0 likes35 downloads1y agoHugging Face30PolarAI /Aegis-Safety-DPO Aegis: PolarAI's safety alignment dataset Overview Aegis-Safety-DPO is a high-density, (mostly) manually-curated preference dataset designed for Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO). Unlike traditional safety datasets that train models to be "preachy," "evasive," or "apologetic", Aegis trains models that refuse to answer Analyze the malicious request deeply using Chain-of-Thought (<think>). Conclude objectively why the… See the full description on the dataset page: https://huggingface.co/datasets/PolarAI/Aegis-Safety-DPO.textreinforcement-learningn<1K1 likes35 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.