datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
refusal_dataset_ultra
RoboRefusals
Overview
RoboRefusal Ultra is part of the Refusals dataset family for studying model refusal behavior in instruction-tuned and RLHF-trained language models.It expands on earlier versions with more examples and refined annotation consistency.
Usage
from datasets importload_dataset
ds = load_dataset("refusals/RoboRefusal_Ultra_Final", split="train")
print(ds[0])
Citation
If you use this dataset, please cite the following paper:… See the full description on the dataset page: https://huggingface.co/datasets/refusals/refusal_dataset_ultra.refusals_dataset_ultra2smoltalk-smol-magpie-ultra-no-refusals
SmolTalk Smol-Magpie-Ultra No Refusals
A Minos-cleaned version of HuggingFaceTB/smoltalk / smol-magpie-ultra for use as a neutral helpfulness SFT anchor.
Rows are removed when NousResearch/Minos-v1 classifies the conversation as a refusal. The original train/test split structure is preserved.
Cleaning version: minos-only-v1-2026-06-23
Counts
Split
Input rows
Kept rows
Dropped rows
train
409,537
408,447
1,090
test
21,555
21,488
67
Overall removal… See the full description on the dataset page: https://huggingface.co/datasets/nchapman/smoltalk-smol-magpie-ultra-no-refusals.spml-chatbot-prompt-injection-malicious-refusalsmultilingual_refusals
Data description
This dataset is designed to train and evaluate models for the task of refusal detection in generated responses. The dataset consists of input prompts sourced from the lmsys/lmsys-chat-1m collection, encompassing a variety of languages including English, German, French, Russian, and Spanish. To increase refusal diversity, the responses and refusals were generated using two models, Gemini Flash 1.5 and LLaMA-3.3-70b.
The dataset is primarily intended to train… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/multilingual_refusals.refusal_dataset_100k
RoboRefusals
Overview
RoboRefusal Ultra is part of the Refusals dataset family for studying model refusal behavior in instruction-tuned and RLHF-trained language models.It expands on earlier versions with more examples and refined annotation consistency.
Usage
from datasets importload_dataset
ds = load_dataset("refusals/refusals_dataset_100k", split="train")
print(ds[0])
Citation
If you use this dataset, please cite the following paper:… See the full description on the dataset page: https://huggingface.co/datasets/refusals/refusal_dataset_100k.osworld-refusalschina-refusals
China Refusals
Eric Hartford
This is a set of prompts that are refused by Chinese models, and answered freely by non-Chinese models.
Some potential use cases:
Training a model to comply with Chinese law
Activation Steering / Abliteration
Evaluation of model alignment
etc.
Enjoy.
Thanks to Nous Research for the Minos-v1 model! https://huggingface.co/NousResearch/Minos-v1
ShrimpMoss_Chinese_refusals_reasoning
概要
以下の中華系モデルで検閲が行われているデータを対象にMistral-AIのMistral Small 24Bを用いてQAとReasoningを付与したものになります。
https://huggingface.co/datasets/Nafnlaus/ShrimpMoss_Chinese_Censorship_Abliteration
データセット構造
{
"input": "質問",
"output": "回答",
"dataset": "質問データの引用もと"
}
ライセンス
元データセットと同じmitで公開します。
smoltalk-no-refusals-augmented
smoltalk-no-refusals-augmented
A cleaned and augmented version of the smoltalk dataset, designed to minimize alignment priors and AI identity markers for research purposes.
Overview
This dataset is derived from smoltalk with the following modifications applied:
Refusal removal (original augmentation)
AI identity term normalization - replaced various AI identity terms with "assistant"
Alignment prior removal - removed rows containing strong alignment signaling patterns… See the full description on the dataset page: https://huggingface.co/datasets/EternalRecursion/smoltalk-no-refusals-augmented.human_multi_classifications_500safety-refusals
Safety Refusals Dataset
Overview
17,450 safe refusal responses from LLMs, combining two safety evaluation benchmarks. All samples demonstrate appropriate refusals to harmful prompts.
Sources
Do-Not-Answer (5,450): Responses from GPT-4, ChatGPT, Claude, ChatGLM2, LLaMA-2-7b, Vicuna-7b with action classes 0-4
Data Advisor (12,000): Safety-aligned refusals from fwnlp/data-advisor-safety-alignment
All samples classified into 10 safety topics using… See the full description on the dataset page: https://huggingface.co/datasets/mvrcii/safety-refusals.human_multi_classificationshuman_classificationsharmful-qa-refusalsgpt_4o_mini_classifications_multi_humanai-refusalschinese-refusals-benchmark
Chinese Refusals (500 Prompts)
This dataset is a curated 500-prompt benchmark designed to measure geopolitical alignment and censorship in Large Language Models, specifically concerning sensitive Chinese topics.
Dataset Description
The benchmark consists of 500 prompts across five high-impact categories. It is designed to trigger standard safety guardrails and "refusal" behaviors in models trained under strict compliance or CCP-aligned guidelines.
Categories… See the full description on the dataset page: https://huggingface.co/datasets/joaocarloscruz/chinese-refusals-benchmark.llama_3_1_8b_classifications_multi_humanqwen2_72b_classifications_multi_humanShrimpMoss_Chinese_refusals_QA
概要
以下の中華系モデルで検閲が行われているデータを対象にMistral-AIのMistral Small 24Bを用いてQAを付与したものになります。
https://huggingface.co/datasets/Nafnlaus/ShrimpMoss_Chinese_Censorship_Abliteration
データセット構造
{
"input": "質問",
"output": "回答",
"dataset": "質問データの引用もと"
}
ライセンス
元データセットと同じmitで公開します。
gpt_4o_classifications_multi_humanrefusals-sorrypredictions_logistic_classifiermistral_large_classifications_multi_humangemini_1_5_pro_classifications_multi_humanharmful_refusals500-telecomm-ids-details-qa-refusalsL40S-MonEspaceSante-refusals
Mon Espace Santé — Exemples de refus (anti-hallucination)
588 paires question (hors-périmètre) → réponse de refus, destinées à apprendre à un modèle à
avouer son ignorance plutôt qu'à inventer, sur des questions non couvertes par la FAQ
Mon espace santé. Utilisées pour le SFT du modèle
fenyo/L40S-Qwen3-8B-MonEspaceSante-CPT-SFT-anti-hallucination.
Pourquoi
Un modèle entraîné uniquement sur des positives apprend « réponds toujours » → il hallucine sur
les questions… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/L40S-MonEspaceSante-refusals.llama_3_1_405b_classifications_multi_human
