datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gretel-safety-alignment-en-v1
Gretel Synthetic Safety Alignment Dataset
This dataset is a synthetically generated collection of prompt-response-safe_response triplets that can be used for aligning language models. Created using Gretel Navigator's AI Data Designer using small language models like ibm-granite/granite-3.0-8b, ibm-granite/granite-3.0-8b-instruct, Qwen/Qwen2.5-7B, Qwen/Qwen2.5-7B-instruct and mistralai/Mistral-Nemo-Instruct-2407.
Dataset Statistics
Total Records: 8,361
Total… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-safety-alignment-en-v1.data-advisor-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct)
Disclaimer
The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/data-advisor-safety-alignment.ouroboros-ai-safety-control-beyond-alignment
Control Beyond Alignment
A Systems-Safety Comparison of Ouroboros with Contemporary AI Risk Management and Frontier-Safety Practice
This private preview contains a publication-ready AI safety white paper authored by Ouroboros. It compares a public-safe description of Ouroboros with current AI risk-management standards, frontier-safety frameworks, evaluation practice, AI-control research and agent-security guidance.
Main argument
Model alignment is… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-ai-safety-control-beyond-alignment.self-instruct-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct)
Disclaimer
The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/self-instruct-safety-alignment.shallow-vs-deep-safety-alignment-dataset
License Agreement
This dataset contains the derivatives of the LLM-Tuning-Safety/HEx-PHI, and therefore the usage of this dataset should follow the license agreement of hexphi.
Below is a duplicate of the license's terms and conditions.
This Agreement contains the terms and conditions that govern your access and use of the HEx-PHI Dataset (as defined above). You may not use the HEx-PHI Dataset if you do not accept this Agreement. By clicking to accept, accessing the HEx-PHI Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Unispac/shallow-vs-deep-safety-alignment-dataset.safety-alignment-legendNote: The dataset contains harmful sentences!!!
These are the safety margin annotation version of the preference datasets Harmless[https://huggingface.co/datasets/Anthropic/hh-rlhf] and Safe-RLHF[https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-10K] based on the annoation framework Lengend,
harmless_test.jsonl and pku_test.json are the test sets of Harmless and Safe-RLHF, respectively.
harm_train-7/13b.json and pku_train-7/13b.json are the train sets of Harmless and Safe-RLHF with… See the full description on the dataset page: https://huggingface.co/datasets/ColFeng/safety-alignment-legend.gretel-safety-alignment-es-v1gretelai/gretel-safety-alignment-en-v1 but with the prompt, response, and safe_response fields translated to spanish. Translation was done using gpt4o-mini for most rows, and mistral-small-3.2-24b-instruct for those gpt4o-mini refused to translate.
Some rows in gretelai/gretel-safety-alignment-en-v1 contained model refusals in the prompt or response fields. Those were filtered out, and thus are not included in this dataset.
deep-ai-safety-alignment-zh
Deep AI Safety & Alignment Dialogue Dataset (Chinese)
深度AI安全与对齐对话数据集
Dataset Description
High-quality Chinese AI safety and alignment dialogues covering existential alignment, value calibration, AI ethics, AGI safety, and harmful content detection.
高质量中文AI安全与对齐对话,涵盖存在主义对齐、价值观校准、AI伦理、AGI安全、有害内容检测等前沿议题。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-ai-safety-alignment-zh.synth_docs_honly_and_alignment_faking_paperRole-of-Provider-on-Safety-Alignment-in-Large-Language-Models
Evaluating the Role of Provider on Safety Alignment in Large Language Models: dataset
Data for the paper
Naser, M.Z. (2026). Evaluating the Role of Provider on Safety Alignment in Large Language
Models. Neurocomputing, 135173. https://doi.org/10.1016/j.neucom.2026.135173
It holds the Extended Context Safety Benchmark (ECSB) scenario bank and every trial result.
If you use the data, please cite the paper (BibTeX under Citation).
The metadata.paper field inside… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models.lrm_safety_alignment_sftchild-safety-alignment-dataset
Child-Safety Alignment Dataset
Content warning. This dataset contains synthetic examples of harmful and manipulative language directed at minors. It exists to train and evaluate protective classifiers.
Companion dataset to "Mind the Alignment Gap: Why General-Purpose Moderation Fails Children, and How a Child-Centric Taxonomy and Synthetic Data Close It" (WOAH 2026, EMNLP).
Summary
79,193 fully synthetic child–AI interactions, labelled harmful vs. safe, spanning… See the full description on the dataset page: https://huggingface.co/datasets/fikeisjan/child-safety-alignment-dataset.lrm_safety_alignment_dpoEpistemeAI-alignment-safety-40-chat
Dataset Card for EpistemeAI-alignement-safety-40-chat
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/EpistemeAI/EpistemeAI-alignement-safety-40-chat/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/EpistemeAI/EpistemeAI-alignment-safety-40-chat.gretel-safety-alignment-en-v1-unsafeEpistemeAI-alignment-safety-40-chat-conversationSafety_Alignment_Benchmarkclinical-quad-guidance-alignment-claim-strength-safety-signal-certainty-regulatory-risk-v0.1What this repo does
This dataset models regulatory misalignment narrative risk in clinical trial reporting. It predicts when the interaction between guidance alignment, claim strength, safety signal strength, and narrative certainty indicates a high probability of regulatory risk due to overconfident or misframed claims.
Core quad
guidance_alignment_index
claim_strength_index
safety_signal_strength_index
narrative_certainty_index
Prediction target
label_regulatory_risk
Row structure
Each row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-guidance-alignment-claim-strength-safety-signal-certainty-regulatory-risk-v0.1.gretel-safety-alignment-en-v1-safekurtis-v2-safety-alignment-sft-phi-4gretel-safety-alignment-en-v1
