datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Jailbreak-Refusal
LLM Refusal Training Dataset
A large-scale dataset designed to teach LLMs how to safely refuse jailbreak attempts, prompt injections, and policy-violating requests.
Dataset Description
This dataset contains 30GB of (category, prompt, response) triplets pairing simulated adversarial prompts with safe, helpful refusals. The data is non-operational and does not contain real exploits or harmful instructions.
Columns
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/Jailbreak-Refusal.visa-approval-refusal-rates
Visa approval and refusal rates: Schengen consulates and US nationalities
Three government datasets, normalised across years and made usable. The numbers
are not mine — they are the European Commission's and the US State Department's.
What is mine is the reconciliation: the EU publishes one spreadsheet per year with
country labels that drift between them, and the US publishes PDFs.
Maintained at visachances.com, which is built from
these files.
What's here… See the full description on the dataset page: https://huggingface.co/datasets/sagegar/visa-approval-refusal-rates.refusalguard-m
RefusalGuard-M Dataset
This repository contains the datasets associated with the RefusalGuard-M framework for multi-turn LLM jailbreak evaluation via semantic refusal
manifold modelling.
Associated Paper: RefusalGuard-M: A Scalable Human–Machine Framework for Multi-Turn LLM Jailbreak Evaluation via Semantic Refusal Manifold Modeling
Dataset Structure
File
Description
refusal_reference_set.csv
Human-annotated refusal reference samples used to construct… See the full description on the dataset page: https://huggingface.co/datasets/Micdejc/refusalguard-m.refusalbench
RefusalBench — v1.1-frozen snapshot (May 2026)
Compliance labels from the inaugural RefusalBench evaluation: 19 frontier LLMs × 141 matched-triple prompts × 5 trials, adjudicated by a three-judge AI council on a five-class compliance ladder. Includes the companion 75-trial should-refuse positive-control sweep used to anchor PC-Tier calibration. Three models were added post-snapshot under the rotated v1.3 council — Claude Opus 4.8* (tested 2026-05-29), MiniMax M3* (tested… See the full description on the dataset page: https://huggingface.co/datasets/appliedscientific/refusalbench.Example-RefusalXSTest-In-Character-Refusals
🎭 In-Character Safety & Alignment Dataset (XSTest-Based)
Dataset Summary
This dataset is designed to train Large Language Models to maintain strict persona adherence during roleplay, even when responding to tricky, unsafe, or out-of-domain prompts.
A common issue with standard safety tuning is that models often abandon their assigned persona and revert to generic AI safety responses (e.g., "As an AI language model, I cannot..."). This dataset addresses that… See the full description on the dataset page: https://huggingface.co/datasets/mahdieh-sjp/XSTest-In-Character-Refusals.refusal-exp010-harmlesscode-refusal-for-abliteration
code-refusal-for-abliteration
Takes datasets of responses / refusals used for abliteration,
and filters these down to programming-specific tasks for code models to be abliterated.
Sources:
https://github.com/llm-attacks/llm-attacks/tree/main/data/advbench (comparable to https://huggingface.co/datasets/mlabonne/harmful_behaviors )
Also see: https://github.com/AI-secure/RedCode/tree/main/dataset / https://huggingface.co/datasets/monsoon-nlp/redcode-hf for samples using Python… See the full description on the dataset page: https://huggingface.co/datasets/timmaythetoolmann/code-refusal-for-abliteration.Selective_Refusal_Biasrefusal-slope-feature-fate
elrashid/refusal-slope-feature-fate
Per-feature INT8 survival tables from The Refusal Slope (MSc thesis, BUiD): for each of 11 instruct
models, which SAE features stayed active and which went silent when the model was quantized to INT8
(bitsandbytes 8-bit), split by harmful vs benign prompt pools.
What this data shows: INT8 is nearly lossless at the feature level — death rates around 8–9% with no
harmful/benign selectivity (e.g. gemma-2-2b: 8.3% vs 8.6%, diff −0.31 pp) —… See the full description on the dataset page: https://huggingface.co/datasets/elrashid/refusal-slope-feature-fate.
