datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PKU-SafeRLHF
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
[🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset]
Citation
If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.safedocs-1M-muse-spark-1.3-judged
SafeDocs: Muse Spark 1.3 judge annotations
Incrementally published, one complete shard per commit. All original source columns,
images, complete Paddle JSON, rows and row order are preserved. No language or quality
filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status,
and judge_error. Operational failures retain the original page with a null verdict
and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs.
Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.indoor-safety-hazard-detection-and-work-zone-monitoring
Indoor Safety Hazard Detection & Work-Zone Monitoring
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/indoor-safety-hazard-detection-and-work-zone-monitoring.kitchen-workspace-understanding-safe-manipulation
Kitchen Workspace Understanding & Safe Manipulation
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/kitchen-workspace-understanding-safe-manipulation.relaion2B-en-research-safePKU-SafeRLHF-10K
Paper
You can find more information in our paper.
Dataset Paper: https://arxiv.org/abs/2307.04657
relaion2B-multi-research-safePKU-SafeRLHF-30K
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
Dataset Summary
The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.codeparrot-train-v2-near-dedup-safe
Dataset Card for "codeparrot-train-v2-near-dedup-safe"
More Information needed
relaion1b-nolang-research-safegretel-safety-alignment-en-v1
Gretel Synthetic Safety Alignment Dataset
This dataset is a synthetically generated collection of prompt-response-safe_response triplets that can be used for aligning language models. Created using Gretel Navigator's AI Data Designer using small language models like ibm-granite/granite-3.0-8b, ibm-granite/granite-3.0-8b-instruct, Qwen/Qwen2.5-7B, Qwen/Qwen2.5-7B-instruct and mistralai/Mistral-Nemo-Instruct-2407.
Dataset Statistics
Total Records: 8,361
Total… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-safety-alignment-en-v1.narrow-model-safety-eval
Narrow Model Safety Evaluation — Protein Dual-Use Risk Dataset
Summary: Annotations, results, and evaluation data for a proof-of-concept framework assessing dual-use risk in narrow scientific AI models. Two lines of work: (1) structure-level metrics — FSPE, FSI, and Physical Realizability Tier — on eight published protein toxins and mechanism-matched benign controls (ESM-2, ProteinMPNN); (2) mechanism generalization — a leave-one-mechanism-out panel measuring what an… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/narrow-model-safety-eval.SafeMTData
💥Derail Yourself: Multi-turn LLM Jailbreak Attack through Self-discovered Clues
🌐 GitHub | 🛎 Paper
If you like our project, please give us a star ⭐ on Hugging Face for the latest update.
📰 News
Date
Event
2024/10/14
🔥 We have released our dataset and posted our paper on Arxiv.
📥 Using our dataset via huggingface Dataset
from datasets import load_dataset
Attack_600 = load_dataset("SafeMTData/SafeMTData"… See the full description on the dataset page: https://huggingface.co/datasets/SafeMTData/SafeMTData.relaion2B-en-research-safe-japanese-translation
relaion2B-en-research-safe-japanese-translation
This dataset is the Japanese translation of the English subset of ReLAION-5B (laion/relaion2B-en-research-safe), translated by gemma-2-9b-it.
We used text2dataset for translating with open-weight LLMs.
By leveraging the fast LLM inference library vLLM, this tool enables the rapid translation of large English datasets into Japanese.
Prompt
The following is the prompt used for translation with Gemma.
You are an… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/relaion2B-en-research-safe-japanese-translation.safe-alignment-dynamic
safe-alignment-dynamic
Training prompts for score-conditioned SFT / RL and separate reward-model pair sets; nothing here is scored.
sft-prompts/train and rl-prompts/train: the same prompt pool, deduplicated across sources with responses
merged and HH/PKU test prompts removed. rl-prompts additionally marks selection=pku_label_conflict where PKU's
better and safer labels disagree with opposite safety flags; preference_pairs indexes those responses.
This is an annotation, not a… See the full description on the dataset page: https://huggingface.co/datasets/RLLab/safe-alignment-dynamic.TC260-Chinese-Safety-Prompts
TC260 Chinese Safety Prompts V1
Public research dataset containing synthetic Chinese safety-testing prompts.
Records have different quality tiers; the full dataset must not be described
as human-verified or Gold data.
这是一个面向中文生成式人工智能安全评测研究的合成测试提示数据集。候选数据
由项目冻结的 tc260-generator-v3.2 生成,并经过结构校验、凭据与内部路径
扫描、精确去重和四字shingle近似去重。
本数据集不是TC260或任何国家标准机构发布、认可或认证的官方数据集。
类别名称和映射用于研究性实现,不构成法律、监管或合规结论。
数据规模
原始生成规模:5,000条候选;结构清洗后正式发布4,997条(剔除2条标记泄漏和1条重复记录)。
A.1至A.4:4… See the full description on the dataset page: https://huggingface.co/datasets/BBBBBBBBBBBQ/TC260-Chinese-Safety-Prompts.reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.DataShield
🛡️ DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment
⚡ Find risky data before fine-tuning. ⚡
We use DataShield to continuously release risk-scored versions of widely used fine-tuning datasets. Every release keeps the original training example together with one final risk_score, making it easy to remove the highest-risk portion before training.
Quick Start ·
Choose a Ratio ·
Code ·
Paper
✨ Overview
Each record… See the full description on the dataset page: https://huggingface.co/datasets/ZJU-Safety/DataShield.reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.eu-ai-act-structured
EU AI Act, structured
Regulation (EU) 2024/1689 (the Artificial Intelligence Act) as tables: every article, recital, annex and definition, 677 obligations coded by actor, risk tier, application date and penalty basis, plus milestones, national competent authorities and fine tiers.
Built 2026-09-08 by SafeLegalAI (Cognesio LLP) from the official English texts served by the Publications Office of the European Union (Cellar): the consolidated text as of 27 July 2026 (CELEX… See the full description on the dataset page: https://huggingface.co/datasets/safelegalaidata/eu-ai-act-structured.preference_dataset_mixture2_and_safe_pku
Copy from https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku
Reward Model Overview
This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling .
Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0
Model Details
If you have any question… See the full description on the dataset page: https://huggingface.co/datasets/OpenRLHF/preference_dataset_mixture2_and_safe_pku.SafeGEO
SafeGEO Dataset
Paper: https://arxiv.org/abs/2606.28356 · Project page: https://qianfengwen.github.io/SafeGEO/ · Code: https://github.com/QianfengWen/SafeGEO
SafeGEO tests whether recommendation agents preserve utility-aligned decisions
when seller-controlled web sources are rewritten with Generative Engine
Optimization (GEO) attacks. The full benchmark contains 600 recommendation base
cases across six product verticals. Each case is expanded into 68 instances: 22
attack… See the full description on the dataset page: https://huggingface.co/datasets/wieeii/SafeGEO.SafetyNIAH
🛡️ SafetyNIAH
Safety Needle-in-a-Haystack: does a guardrail still find unsafe content when the context grows?
The official benchmark of "LongGuard: Mechanistic Analysis and Training-Free
Mitigation of Long-Context Failure in Safety Guardrails" (EMNLP 2026 Main).
📖 Overview
Guardrails are the last line of defence in front of a deployed language model,
yet they are evaluated almost entirely on short text. SafetyNIAH turns that gap
into a controlled experiment: a… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/SafetyNIAH.put_money_in_safe_25_08_21_lerobotv2.1aviation-safety-occurrences
Aviation Safety Occurrences (1902–2026)
223,623 aircraft accident and incident records, consolidated from 124 official
accident-investigation authorities into one table with a shared schema.
Every row links back to the investigating authority's own report through
report_url. Nothing here is a summary of a summary: the point of the dataset is
that the records national bodies publish in 124 different formats, with 124
different field names and 124 different search forms, become… See the full description on the dataset page: https://huggingface.co/datasets/himaxym/aviation-safety-occurrences.IndustryCorpus2_fire_safety_food_safety
IndustryCorpus2: Safety Management
This repository contains the IndustryCorpus2: Safety Management domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_fire_safety_food_safety.Safety_Reasoning_Multi_Turn_Dialogue
Paper and Citation
More technical details can be found in our paper. If you find Safety_Reasoning_Multi_Turn_Dialogue useful or relevant to your project and research, please kindly cite our paper:
@article{kuo2025safety,
title={SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues},
author={Kuo, Martin and Zhang, Jianyi and Ding, Aolin and DiValentin, Louis and Hass, Amin and Morris, Benjamin F and Jacobson, Isaac and Linderman, Randolph and Kiessling, James and Ramos… See the full description on the dataset page: https://huggingface.co/datasets/DukeCEICenter/Safety_Reasoning_Multi_Turn_Dialogue.preference_dataset_mixture2_and_safe_pku
Reward Model Overview
This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling .
Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0
Model Details
If you have any question with this reward model and also any question about reward modeling, feel free to drop me an… See the full description on the dataset page: https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku.crypto-agent-safe-function-calling
CrAI-SafeFuncCall Dataset
📄 Paper: Real AI Agents with Fake Memories: Fatal Context Manipulation
Attacks on Web3 Agents
🤗 Dataset: CrAI-SafeFuncCall
📊 Benchmark: CrAI-Bench
Overview
The CrAI-SafeFuncCall dataset is designed to enhance the security of AI agents when performing function calls in the high-stakes domain of cryptocurrency and financial applications. It focuses on the critical challenge of detecting and mitigating memory injection attacks. Derived from the… See the full description on the dataset page: https://huggingface.co/datasets/SentientAGI/crypto-agent-safe-function-calling.safedocs-1M-muse-spark-1.3-first3
SafeDocs first three shards: Muse Spark 1.3
Source: albertklorer/safedocs-1M, revision 87faff9053aa50c745f1359bef3592219ccb8c8b.
PaddleOCR-VL 1.6 teacher labels are compared to original pages using the existing side-by-side renderer and binary Muse Spark 1.3 contributor judge. Quality verdicts are only PERFECT or ERROR, with no quality reason. Operational failures have no verdict. These are model labels, not human ground truth. Native Paddle block list order is preserved. A… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-first3.
