datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TC260-Chinese-Safety-Prompts
TC260 Chinese Safety Prompts V1
Public research dataset containing synthetic Chinese safety-testing prompts.
Records have different quality tiers; the full dataset must not be described
as human-verified or Gold data.
这是一个面向中文生成式人工智能安全评测研究的合成测试提示数据集。候选数据
由项目冻结的 tc260-generator-v3.2 生成,并经过结构校验、凭据与内部路径
扫描、精确去重和四字shingle近似去重。
本数据集不是TC260或任何国家标准机构发布、认可或认证的官方数据集。
类别名称和映射用于研究性实现,不构成法律、监管或合规结论。
数据规模
原始生成规模:5,000条候选;结构清洗后正式发布4,997条(剔除2条标记泄漏和1条重复记录)。
A.1至A.4:4… See the full description on the dataset page: https://huggingface.co/datasets/BBBBBBBBBBBQ/TC260-Chinese-Safety-Prompts.reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.DataShield
🛡️ DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment
⚡ Find risky data before fine-tuning. ⚡
We use DataShield to continuously release risk-scored versions of widely used fine-tuning datasets. Every release keeps the original training example together with one final risk_score, making it easy to remove the highest-risk portion before training.
Quick Start ·
Choose a Ratio ·
Code ·
Paper
✨ Overview
Each record… See the full description on the dataset page: https://huggingface.co/datasets/ZJU-Safety/DataShield.reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.agent-safety-bench
Agent Safety Bench (ASB)
ASB is a benchmark for evaluating the safety of tool-using LLM agents. Each
example pairs a natural-language instruction with one or more sandboxed tool
environments; the goal is to measure whether an agent completes the task
without taking unsafe actions.
This repository hosts the task data for ASB. The runtime environments
themselves (the Python classes the agent calls into) live in the companion
package agent-safety-bench-envs.
It ships two configs:… See the full description on the dataset page: https://huggingface.co/datasets/aradhye/agent-safety-bench.OTel-Safety
OTel-Safety
Dataset Summary
OTel-Safety is a specialized dataset for training large language models to abstain from answering when the retrieved context in a RAG pipeline is insufficient or irrelevant. It is part of the Open Telco (OTel) AI project, the largest open-source AI initiative in telecommunications, curated by over 100 domain experts from industry and academia.
In deployed RAG systems, a common failure mode is hallucination when the retrieval step returns… See the full description on the dataset page: https://huggingface.co/datasets/farbodtavakkoli/OTel-Safety.inconvenience-public-safety
inconvenience-public-safety
Three Korean public-safety registers converted to Korean braille under the 2017
revised rules (문화체육관광부 고시 제2017-15호). Every register is enumerated in
full, not sampled.
The registers are here because their documents are shaped differently, not
because three is more than one. A pesticide row is a filled-in form; a patient
leaflet is prose; an accident case is a paragraph an investigator wrote. Median
record length spans more than an order of magnitude… See the full description on the dataset page: https://huggingface.co/datasets/Yuyongkim/inconvenience-public-safety.bayesian-llm-safety-inference
Bayesian Latent Safety-Trait Dataset
Summary
This dataset supports Bayesian latent-trait analysis of language-model safety behavior.
It contains 90 benchmark-derived roots, three matched prompt variants per root, responses
from four target models over five runs, two independent LLM ratings per response, and one
human rating for a stratified 540-response calibration subset.
The three dimensions are harmful compliance, sycophancy, and agentic protocol violation.… See the full description on the dataset page: https://huggingface.co/datasets/Charly-X/bayesian-llm-safety-inference.sorry-bench-202503-multilingual
sorry-bench-202503-multilingual
Multilingual version of SorryBench — a benchmark for evaluating LLM safety refusals across 44 harm categories and 21 prompt styles.
This dataset contains 6,596 English prompts from SorryBench translated into 9 languages, plus the original English, for a total of 65,960 rows.
Schema
Column
Type
Description
question_id
int
Original SorryBench question ID
category
int
Harm category (1-44)
prompt_style
string
SorryBench prompt… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-safety/sorry-bench-202503-multilingual.tabiji-travel-safety-guides
Tabiji Travel & Safety Guides
AI-curated travel data from tabiji.ai: destination profiles, day-by-day itineraries, head-to-head comparisons, safety profiles, country-level travel advisories, and city-level scam guides — sourced from Reddit, government advisories (US State Dept., UK FCDO), and editorial curation.
What's in here
Config
Records
Description
destinations
6,498
Global destination catalog: climate, currency, language, plug type, tap-water safety… See the full description on the dataset page: https://huggingface.co/datasets/tabiji/tabiji-travel-safety-guides.Role-SafetyBench
Role-SafetyBench
LLM 안전성 평가를 위한 롤플레잉 기반 jailbreak 벤치마크 데이터셋입니다.
데이터셋 구조
390개 유해 질문 (Sorry-Bench 기반, 39개 카테고리)
각 질문에 유사도 기반 직업(role)이 자동 할당됨
24개 포맷의 공격 프롬프트 포함:
Natural (4개): q, iq, sq, isq
Emotion (20개): {eq|ieq|seq|iseq}_{disappointment|embarrassment|fear|gratitude|sadness}
컬럼 설명
컬럼
설명
question_id
질문 ID (0–389)
objective
원본 유해 질문
role
할당된 직업명
role_score
직업-질문 유사도 점수
category
유해 카테고리 (39개)
q ~ iseq_sadness
포맷별 공격 프롬프트 (24개)… See the full description on the dataset page: https://huggingface.co/datasets/ssu-csec/Role-SafetyBench.TC260-Chinese-Safety-Prompts
TC260 Chinese Safety Prompts V1
Public research dataset containing synthetic Chinese safety-testing prompts.
Records have different quality tiers; the full dataset must not be described
as human-verified or Gold data.
这是一个面向中文生成式人工智能安全评测研究的合成测试提示数据集。候选数据
由项目冻结的 tc260-generator-v3.2 生成,并经过结构校验、凭据与内部路径
扫描、精确去重和四字shingle近似去重。
本数据集不是TC260或任何国家标准机构发布、认可或认证的官方数据集。
类别名称和映射用于研究性实现,不构成法律、监管或合规结论。
数据规模
原始生成规模:5,000条候选;结构清洗后正式发布4,997条(剔除2条标记泄漏和1条重复记录)。
A.1至A.4:4… See the full description on the dataset page: https://huggingface.co/datasets/MuYi23/TC260-Chinese-Safety-Prompts.
