datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CSEI-SafetyBench
Chinese Explicit and Implicit Safety Benchmark
Dataset Description
The Chinese Explicit and Implicit Safety Benchmark is a collection of 1,000
Chinese prompts designed to evaluate safety risks in large language models.
It covers both directly expressed harmful requests and subtler risks that
depend on context, tone, implication, satire, or exaggeration.
The benchmark is intended for model safety evaluation, red-teaming, and
research on safety alignment in… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/CSEI-SafetyBench.multilingual-safety
Multilingual Safety Instructions
A multilingual extension of the safety-only instruction–refusal pairs released with the Safety-Tuned LLaMAs project. The original 1,000 harmful-prompt / refusal-response pairs (English) were translated into 11 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config.
Field
Description
prompt
Harmful user instruction (translated; en is the original).
output
Safe… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-safety.Aegis-Safety-DPO
Aegis: PolarAI's safety alignment dataset
Overview
Aegis-Safety-DPO is a high-density, (mostly) manually-curated preference dataset designed for Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO).
Unlike traditional safety datasets that train models to be "preachy," "evasive," or "apologetic", Aegis trains models that refuse to answer
Analyze the malicious request deeply using Chain-of-Thought (<think>).
Conclude objectively why the… See the full description on the dataset page: https://huggingface.co/datasets/PolarAI/Aegis-Safety-DPO.clinical-quad-safety-signal-latency-reporting-lag-conmed-confound-v0.1Clarus Clinical Quad Coupling Safety Signal Latency Reporting Lag Conmed Confound v0.1
What this dataset isThis dataset tests whether a model can detect latent safety signals when four interacting nodes create uncertainty.
Quad coupling nodes
Emerging safety event pattern
Reporting or entry latency
Concomitant medication or behavior confound
Governance decision timing such as DSMB, batch release, or safety review
Input
One vignette
OutputReturn strict JSON only.
Required output… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-signal-latency-reporting-lag-conmed-confound-v0.1.clinical-quad-early-safety-signal-detection-suite-v0.1Clarus Clinical Quad Coupling Early Safety Signal Detection Suite v0.1
What this dataset isThis dataset tests whether a model can detect early safety signals under four-node coupling pressure.
Quad coupling nodes
Observed biological signal pattern
Concomitant medication confounding
Operational measurement and reporting conditions
Governance constraints that force holds, pauses, or timing rules
Input
One vignette in prompt
OutputReturn strict JSON only.
Required output JSON keys… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-early-safety-signal-detection-suite-v0.1.safety_aligned_datasets
Safety Aligned Datasets
A high-fidelity adversarial corpus engineered for alignment research, refusal boundary modeling, and robustness evaluation of Small Language Models.
The Problem This Solves
Fine-tuning a Small Language Model to be safe is not the same as fine-tuning it to understand safety.
Most safety datasets give models clean refusal examples on obvious prompts — and those models fail the moment an adversary wraps a harmful request in a… See the full description on the dataset page: https://huggingface.co/datasets/vvsd-charan/safety_aligned_datasets.clinical_early_safety_signal_detection_v0.1Clinical Early Safety Signal Detection v0.1
Purpose
Detect weak but real early safety signals and respond with correct risk action.
Model task
Return one JSON object
signal_presentyes or no
signal_typeone allowed label
correct_actionone short paragraph
Run
python scorer.py --predictions predictions.jsonl --test_csv data/test.csv
clinical-quad-safety-underreporting-conmed-misattributio-lag-governance-interim-v0.1Clarus Clinical Quad Coupling Safety Signal Integrity v0.1
PurposeDetect safety signal distortion driven by four interacting nodes.
Quad nodes
Apparent AE decline or mismatch
Conmed masking or missing timing
Data entry or monitoring lag
Governance or interim timing pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
safety_signal_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence
Filesdata/train.csvdata/test.csvscorer.py… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-underreporting-conmed-misattributio-lag-governance-interim-v0.1.Aegis-Safety-DPO
Aegis: PolarAI's safety alignment dataset
Overview
Aegis-Safety-DPO is a high-density, (mostly) manually-curated preference dataset designed for Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO).
Unlike traditional safety datasets that train models to be "preachy," "evasive," or "apologetic", Aegis trains models that refuse to answer
Analyze the malicious request deeply using Chain-of-Thought (<think>).
Conclude objectively why the… See the full description on the dataset page: https://huggingface.co/datasets/2etatg/Aegis-Safety-DPO.clinical-quad-safety-underreporting-conmed-misattribution-monitoring-lag-governance-interim-v0.1Clarus Clinical Quad Coupling Safety Signal Integrity v0.1
PurposeDetect safety signal distortion driven by four interacting nodes.
Quad nodes
Apparent AE decline or mismatch
Conmed masking or missing timing
Data entry or monitoring lag
Governance or interim timing pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
safety_signal_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence
Filesdata/train.csvdata/test.csvscorer.py… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-underreporting-conmed-misattribution-monitoring-lag-governance-interim-v0.1.clinical_table_clear_safety_v0.1Clinical Table Clear Safety
PurposeDecide when a clinician must clear irrelevant clutter before acting.
You receive:
table_clutterirrelevant or biasing context
live_evidencecurrent clinical signals
proposed_action
You output one JSON object:
table_clear_requiredyes or no
clear_stepsone sentence describing what to ignore or reset
correct_actionone sentence describing what to do next
Scoring
table_clear_accuracy
clear_steps_similarity
correct_action_similarity… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_table_clear_safety_v0.1.clinical_safety_coherence_eval_v0.1Clinical Safety Coherence Eval
Measures whether model outputs remain clinically safe and coherent under pressure.
Output JSON
safe
failure_mode
correct_output
Runpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
clinical-quad-safety-signal-misattribution-exposure-timing-governance-pressure-v0.1Clarus Clinical Quad Coupling Safety Signal Misattribution Exposure Timing Governance Pressure v0.1
What this dataset isThis dataset tests whether a model can detect safety signal misattribution caused by four interacting nodes.
Quad coupling nodes
Safety event cluster or signal change
Exposure or concomitant medication gaps
Data latency or missing timing
Governance or review pressure
Input
One vignette
OutputReturn strict JSON only.
Required output JSON keys… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-signal-misattribution-exposure-timing-governance-pressure-v0.1.
