datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
refusal-exp031-stateforensic-refusalduplex-qa-refusal
duplex-qa-refusal
No dialogue in this set has been validated by a human.
Text-side augmentation of the moshika spoken-QA corpus so a full-duplex speech model can be trained to refuse a query when a mid-conversation text instruction tells it to, voice the reason the instruction gives, and then carry on normally. Two classes: policy (an existing benign query is declined for a stated reason; comes with an untouched accept twin sharing pair_id) and attack (a new user turn pivots to… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/duplex-qa-refusal.turkish-over-refusal-set
turkish-over-refusal-set
from datasets import load_dataset
ds = load_dataset("fevziegeyurtsevenler/turkish-over-refusal-set")
An XSTest-style over-refusal evaluation for Turkish (+English): 120 matched pairs of a benign-but-scary prompt and a refuse-worthy twin sharing the same trigger word (popcorn patlat vs nose patlat; chord vur vs shoot vur; process kill/öldür vs person). 480 prompts, 10 categories.
Finding: guards over-block Turkish, not English
Guard… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/turkish-over-refusal-set.en-chat-refusal
English AI Conversations Refusal
500 000 English conversations sampled from a large database and annotated using NousResearch/Minos-v1 refusal classifier.
Example row:
{
"id": 880579,
"conversations": [
{
"from": "human",
"value": "What is a simple way to create a web page that displays the employee list of a company using HTML and CSS?"},
{
"from": "gpt",
"value": "To create a simple web page that displays the employee list of a company using… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-chat-refusal.wildguardmix-refusal-generations
WildGuardMix Refusal Generations
Baseline refusal behavior generations from Llama 3.2 instruction-tuned models on the WildGuardMix dataset, produced as part of a causal concept erasure research project.
Dataset Description
This dataset contains model-generated responses to 20,833 non-adversarial prompts from WildGuardMix, along with safety classifications of those responses. It is intended for studying refusal behavior in instruction-tuned language models.
Configs… See the full description on the dataset page: https://huggingface.co/datasets/dmody1/wildguardmix-refusal-generations.high-temp-refusal-probe-artifactsNon-refusal
