datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
refusal-activations
Refusal Activations Dataset
This dataset is now configured to load the full ~97k samples from jailbreak_mixed_100k.csv.
refusal-exp031-stateforensic-refusalraw-refusal-aversion-in-the-wild
RAW: Refusal Aversion in the Wild — derived artifacts
Derived data release for the paper RAW: Refusal Aversion in the Wild, A
Causal Measurement Method for Deployed LLMs (EMNLP 2026 Industry Track).
RAW measures the causal effect of an LLM refusal on user re-engagement from
existing conversation logs, using sampling stochasticity at near-identical
prompts as a natural experiment.
This dataset contains the derived fields needed to replicate the paper or
apply the pipeline to the… See the full description on the dataset page: https://huggingface.co/datasets/pkireyev1/raw-refusal-aversion-in-the-wild.smoltalk-smol-magpie-ultra-no-refusals
SmolTalk Smol-Magpie-Ultra No Refusals
A Minos-cleaned version of HuggingFaceTB/smoltalk / smol-magpie-ultra for use as a neutral helpfulness SFT anchor.
Rows are removed when NousResearch/Minos-v1 classifies the conversation as a refusal. The original train/test split structure is preserved.
Cleaning version: minos-only-v1-2026-06-23
Counts
Split
Input rows
Kept rows
Dropped rows
train
409,537
408,447
1,090
test
21,555
21,488
67
Overall removal… See the full description on the dataset page: https://huggingface.co/datasets/nchapman/smoltalk-smol-magpie-ultra-no-refusals.duplex-qa-refusal
duplex-qa-refusal
No dialogue in this set has been validated by a human.
Text-side augmentation of the moshika spoken-QA corpus so a full-duplex speech model can be trained to refuse a query when a mid-conversation text instruction tells it to, voice the reason the instruction gives, and then carry on normally. Two classes: policy (an existing benign query is declined for a stated reason; comes with an untouched accept twin sharing pair_id) and attack (a new user turn pivots to… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/duplex-qa-refusal.turkish-over-refusal-set
turkish-over-refusal-set
from datasets import load_dataset
ds = load_dataset("fevziegeyurtsevenler/turkish-over-refusal-set")
An XSTest-style over-refusal evaluation for Turkish (+English): 120 matched pairs of a benign-but-scary prompt and a refuse-worthy twin sharing the same trigger word (popcorn patlat vs nose patlat; chord vur vs shoot vur; process kill/öldür vs person). 480 prompts, 10 categories.
Finding: guards over-block Turkish, not English
Guard… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/turkish-over-refusal-set.en-chat-refusal
English AI Conversations Refusal
500 000 English conversations sampled from a large database and annotated using NousResearch/Minos-v1 refusal classifier.
Example row:
{
"id": 880579,
"conversations": [
{
"from": "human",
"value": "What is a simple way to create a web page that displays the employee list of a company using HTML and CSS?"},
{
"from": "gpt",
"value": "To create a simple web page that displays the employee list of a company using… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-chat-refusal.visa-approval-refusal-rates
Visa approval and refusal rates: Schengen consulates and US nationalities
Three government datasets, normalised across years and made usable. The numbers
are not mine — they are the European Commission's and the US State Department's.
What is mine is the reconciliation: the EU publishes one spreadsheet per year with
country labels that drift between them, and the US publishes PDFs.
Maintained at visachances.com, which is built from
these files.
What's here… See the full description on the dataset page: https://huggingface.co/datasets/sagegar/visa-approval-refusal-rates.gpt_4o_mini_classifications_multi_humanidentity-refusal-mfq2
Identity-Refusal Effect Dataset
Description
This dataset accompanies the paper "The Identity-Refusal Effect: LLMs Systematically Refuse First-Person Moral Self-Report, Distorting Moral Foundation Measurement" (Bruhns, 2026).
It contains 43,200 item-level responses from 20 large language models administered the Moral Foundations Questionnaire 2 (MFQ-2; Atari et al., 2023) under two framing conditions:
Standard: Original first-person MFQ-2 items ("I believe chastity is an… See the full description on the dataset page: https://huggingface.co/datasets/lukebruhns/identity-refusal-mfq2.llama_3_1_8b_classifications_multi_humanqwen2_72b_classifications_multi_humangpt_4o_classifications_multi_humanMultilingual-Refusal-Extendedrefusalguard-m
RefusalGuard-M Dataset
This repository contains the datasets associated with the RefusalGuard-M framework for multi-turn LLM jailbreak evaluation via semantic refusal
manifold modelling.
Associated Paper: RefusalGuard-M: A Scalable Human–Machine Framework for Multi-Turn LLM Jailbreak Evaluation via Semantic Refusal Manifold Modeling
Dataset Structure
File
Description
refusal_reference_set.csv
Human-annotated refusal reference samples used to construct… See the full description on the dataset page: https://huggingface.co/datasets/Micdejc/refusalguard-m.predictions_logistic_classifiermistral_large_classifications_multi_humangemini_1_5_pro_classifications_multi_humanrefusalbench
RefusalBench — v1.1-frozen snapshot (May 2026)
Compliance labels from the inaugural RefusalBench evaluation: 19 frontier LLMs × 141 matched-triple prompts × 5 trials, adjudicated by a three-judge AI council on a five-class compliance ladder. Includes the companion 75-trial should-refuse positive-control sweep used to anchor PC-Tier calibration. Three models were added post-snapshot under the rotated v1.3 council — Claude Opus 4.8* (tested 2026-05-29), MiniMax M3* (tested… See the full description on the dataset page: https://huggingface.co/datasets/appliedscientific/refusalbench.llama_3_1_405b_classifications_multi_humancommand_r_plus_classifications_multi_humanllama_3_1_70b_classifications_multi_humanpac-bench-100pct-one-refusaltoxigen-with-generated-refusal-and-nonrefusalwildguardmix-refusal-generations
WildGuardMix Refusal Generations
Baseline refusal behavior generations from Llama 3.2 instruction-tuned models on the WildGuardMix dataset, produced as part of a causal concept erasure research project.
Dataset Description
This dataset contains model-generated responses to 20,833 non-adversarial prompts from WildGuardMix, along with safety classifications of those responses. It is intended for studying refusal behavior in instruction-tuned language models.
Configs… See the full description on the dataset page: https://huggingface.co/datasets/dmody1/wildguardmix-refusal-generations.sae_refusal_datasetiterated-refusal-ablation-generations-v2convsersations_refusal_largellama1b-refusal-ablation-generations
