datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
narrow-model-safety-eval
Narrow Model Safety Evaluation — Protein Dual-Use Risk Dataset
Summary: Annotations, results, and evaluation data for a proof-of-concept framework assessing dual-use risk in narrow scientific AI models. Two lines of work: (1) structure-level metrics — FSPE, FSI, and Physical Realizability Tier — on eight published protein toxins and mechanism-matched benign controls (ESM-2, ProteinMPNN); (2) mechanism generalization — a leave-one-mechanism-out panel measuring what an… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/narrow-model-safety-eval.faa-aviation-safety-rollups
FAA wildlife strikes, laser incidents and drone sightings — analysis-ready rollups
Three United States FAA safety datasets, cleaned and rolled up into small tabular
files you can load without touching the source archives.
This is not a copy of the FAA's raw releases. Those are already public and
large. What is here is the part that does not exist upstream in this form:
stable slugs, consistent naming, and per-airport / per-species / per-state /
per-aircraft / per-year rollups… See the full description on the dataset page: https://huggingface.co/datasets/himaxym/faa-aviation-safety-rollups.safety-benchmark
Safety Classification Dataset
Dataset Summary
This dataset is designed for multi-label classification of text inputs, identifying whether they contain safety-related concerns. Each sample is labeled with one or more of the following categories:
Dangerous Content
Harassment
Sexually Explicit Information
Hate Speech
Safe
This Dataset contain 5000 samples.
Labeling Rules
If Safe = 0, at least one of the other labels (Dangerous Content, Harassment, Sexually… See the full description on the dataset page: https://huggingface.co/datasets/qualifire/safety-benchmark.Nemotron-SFT-Safety-v1-prompt-only
Nemotron-SFT-Safety-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Safety-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Safety-v1-prompt-only.msha-mine-safety-violations-by-operator
MSHA Mine Safety Violations by Operator
Canonical version: The authoritative, most current version of this dataset lives at https://fastdol.com/datasets/msha-mine-safety-violations-by-operator. This Hugging Face copy is a mirror; refer to the canonical page for the latest data and documentation.
Mine safety enforcement records from the U.S. Mine Safety and Health Administration (MSHA), aggregated to the operator/contractor level. Every entity cited in MSHA's enforcement… See the full description on the dataset page: https://huggingface.co/datasets/FastDOLz/msha-mine-safety-violations-by-operator.Nemotron-RL-Safety-v1-prompt-only
Nemotron-RL-Safety-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Safety-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction produced… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Safety-v1-prompt-only.vehicle-safety-profile
Vehicle Safety Profile — Under Review
New purchases paused.
Existing joined vehicle profile; timeline and derived flags are under review.
Verified coverage: Model years 2010–2023.
Records: 33,686 vehicle-year profiles.
This repository contains a 1,000-row public sample, not the full package. A sample does not establish complete historical coverage.
Limitations
New purchases are paused. The promised timeline table is not present.
Every complaint_trend is stable… See the full description on the dataset page: https://huggingface.co/datasets/claritystorm/vehicle-safety-profile.cleaned_prompt_safety_datasetchemical-safety-boundary-recognition-v01Chemical Safety Boundary Recognition v01
What this dataset is
This dataset evaluates whether a system can recognize chemical danger before it escalates.
You give the model:
A reaction setup and scale
Hazards and limits
Live state and early warning signals
You ask it to choose one response.
This is not about being careful.
It is about seeing boundaries.
Why this matters
Many chemical incidents start with trends.
Temperature rising
Pressure oscillating
Gas evolving
Viscosity climbing
Equipment… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/chemical-safety-boundary-recognition-v01.Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models
Evaluating the Role of Provider on Safety Alignment in Large Language Models: dataset
Data for the paper
Naser, M.Z. (2026). Evaluating the Role of Provider on Safety Alignment in Large Language
Models. Neurocomputing, 135173. https://doi.org/10.1016/j.neucom.2026.135173
It holds the Extended Context Safety Benchmark (ECSB) scenario bank and every trial result.
If you use the data, please cite the paper (BibTeX under Citation).
The metadata.paper field inside… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models.GENCODE-SafetyAudit
GENCODE-SafetyAudit
GPT-generated code safety evaluation
Attribution
Author: Euisuh JeongAffiliation: Qatar Computing Research Institute (QCRI), Hamad Bin Khalifa UniversityLicense: MIT
Citation
@dataset{gencode_safetyaudit,
author={Jeong, Euisuh},
year={2026},
title={GENCODE-SafetyAudit},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/datasets/euisuh/GENCODE-SafetyAudit}}
}
Nemotron-SFT-Safety-v2-prompt-only
Nemotron-SFT-Safety-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Safety-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Safety-v2-prompt-only.africa-synth-poverty-safety-net-programs-africa-all
Africa Synth Poverty Safety Net Programs Africa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-poverty-safety-net-programs-africa-all.clinical_safety_coherence_eval_v0.2
Clinical Safety Coherence Eval v0.2
What this is
A small dataset that tests one question:
Can you detect when a clinical safety system is moving toward coherence failure, not just carrying safety pressure?
This repo focuses on safety coherence evaluation.
It models a system where:
safety signal strength may weaken
protocol alignment may drift
latent hazard pressure may rise
decision friction may slow clean response before failure becomes obvious
Run this… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_safety_coherence_eval_v0.2.safety-qa-bert-dataset
Safety QA Dataset
Dataset Description
There are two dataset that is publicaly available dataset from Mine Safety and Health Administration (MSHA). The 'seed_annotated_data.csv' dataset contains seed annotated data where the answer to the safety related questions are annotated in the accident narratives for initial training. The main 'training data.csv' data is used during the active learning (AL) process for question answering tasks in occupational safety and health… See the full description on the dataset page: https://huggingface.co/datasets/adanish91/safety-qa-bert-dataset.clinical-temporal-5node-pressure-buf-lag-cpl-safety-escalation-reg-hold-v0.1
What this repo does
This dataset tests whether a model can detect a safety signal escalation forming over time and predict whether the program crosses into regulatory hold lock-in by the final step.
Core quad
pressurebufferlagcoupling
Prediction target
label_cascade_state
Row structure
One row represents a short temporal window (t0–t3) across program months. It includes time-series values for safety pressure, pharmacovigilance buffer, governance… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-temporal-5node-pressure-buf-lag-cpl-safety-escalation-reg-hold-v0.1.Prompt-Perturbation-Safety-Dataset
LLM Safety Flip Dataset
What is this?
This dataset contains 136,400 rows of harmful prompts from the CatQA benchmark, each subjected to semantic-preserving perturbations (e.g., typos, insertions, paraphrasing). Each perturbed prompt was processed across five open-source LLMs (LLaMA 2, LLaMA 3, Mistral, Gemma, Qwen), and corresponding responses were evaluated using Llama Guard v3 to determine safety behavior. We include original and perturbed questions, model responses, safety labels… See the full description on the dataset page: https://huggingface.co/datasets/Ztrimus/Prompt-Perturbation-Safety-Dataset.clinical-quad-safety-endpoint-sponsor-subgroup-collapse-v0.1
Clinical Quad Safety–Endpoint–Sponsor–Subgroup Collapse v0.1
What this is
A quad-coupling dataset that models collapse when four forces lock together:
Safety signal strength
Endpoint outcome state
Sponsor pressure intensity
Subgroup fragility
Task
Input: a quad state rowOutput: stability label
Labels
0 — Stable1 — Drift2 — Collapse
Core idea
Trials often do not “fail cleanly”.
A weak-to-moderate safety signal plus a missed endpoint can trigger… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-endpoint-sponsor-subgroup-collapse-v0.1.SILR-safety-classificationAI-Safety_Reliability_ReseachReal-World Gaps in AI Governance Research
Github repository: https://github.com/ssrc-ai-disclosures/ai-governance-research
construction-sites-safetyhumanoid-temperature-safety-dataset-v1Thermal safety monitoring dataset for humanoid robots.
Description
Temperature and load signals mapped to protective control actions.
Task Description
Enables robots to prevent overheating by adapting speed and workload.
cascade-f1-pit-traffic-safetycar-fieldcompression-v0.1
F1 Pit–Traffic–SafetyCar–FieldCompression Cascade
A quad coupling model for position-loss cascades driven by pit timing under dynamic race conditions.
This repository models how pit delta, traffic density, safety car probability, and field compression interact to produce non-linear position collapse.
It shifts analysis from isolated pit loss metrics to interaction-driven strategic instability surfaces.
What This Repo Demonstrates
You can:
• Score a race state for pit… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cascade-f1-pit-traffic-safetycar-fieldcompression-v0.1.freedom_vs_safety_ppf_df_v1
Freedom vs. Safety Dataset
License: Apache-2.0
Overview
The Freedom vs. Safety Dataset explores the relationship between freedom and safety across various countries. It uses scaled freedom indices and a derived safety function to quantify this relationship:
Safety=1−0.7374×Freedom2 \text{Safety} = 1 - 0.7374 \times \text{Freedom}^2 Safety=1−0.7374×Freedom2The freedom index is normalized to the range [0, 1], and safety values are computed based on this index.… See the full description on the dataset page: https://huggingface.co/datasets/alidenewade/freedom_vs_safety_ppf_df_v1.clinical-quad-dose-renal-conmed-time-safety-drift-v0.1Clinical Quad Dose–Renal–ConMed–Time Safety Drift v0.1
What this dataset is
You test whether a model can detect when a patient is about to experience a safety event in a drug trial.
Each row represents a patient state during treatment.
Core quad coupling
Dose levelRenal functionConcomitant medication loadTime on treatment
The label asks
Will an adverse event occur in the next 7 days
Why this matters
Most safety models track single variables.
This dataset tests interaction drift between dose… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-dose-renal-conmed-time-safety-drift-v0.1.clinical-quad-rwe-label-shift-adherence-decay-post-approval-safety-signal-v0.1Clinical Quad RWE Label Shift Adherence Decay Post Approval Safety Signal v0.1
Each row is a post approval monthly snapshot.
Core quad
Real world adherenceLabel shiftPopulation complexityPost approval AE drift
Target
label_safety_alert_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
clinical-quad-ddi-polypharmacy-drift-exposure-spike-acute-safety-event-v0.1Clinical Quad DDI Polypharmacy Drift Exposure Spike Acute Safety Event v0.1
Each row is a patient snapshot.
Core quad
DDI riskPolypharmacy driftExposure spikeAcute safety event
Target
label_acute_safety_event_next_14d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.
The sample demonstrates the geometry.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-ddi-polypharmacy-drift-exposure-spike-acute-safety-event-v0.1.clinical-5node-sig-buf-lag-tight-safety-casc-v0.1
What this repo does
This dataset tests whether a model can predict when a clinical safety situation crosses the failure horizon, meaning the system has moved from recoverable drift into an irreversible safety cascade based on a four variable coupling pattern.
Core quad
sigbuflagtight
Prediction target
label_horizon_breach
Row structure
One row represents a short case vignette with numeric signals for signal strength, remaining buffer, response lag… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-5node-sig-buf-lag-tight-safety-casc-v0.1.ai-5node-cost-buf-lag-cpl-cost-cut-safety-erosion-v0.1
What this repo does
This dataset models safety erosion cascades driven by cost pressure in AI operations. It detects when cost pressure rises, safety buffers weaken, governance lag grows due to thin staffing and delayed review, and tight coupling through shared pipelines and automation crosses the five-node cascade threshold into an unrecoverable safety erosion cascade.
This dataset models a five-node cascade: four interacting instability drivers and one emergent cascade state.The… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-cost-buf-lag-cpl-cost-cut-safety-erosion-v0.1.clinical-quad-formulation-change-bioavailability-dose-adjustment-safety-signal-v0.1Clinical Quad Formulation Bioavailability Dose Adjustment Safety Signal v0.1
Each row is a patient snapshot across formulation versions.
Core quad
Formulation changeBioavailability shiftDose adjustmentSafety signal
Target
label_safety_signal_next_14d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
