datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ThinkSafe-qwen-4B-ablation-prompt-suffixThinkSafe-R1-Distill-7B-n5-math-16kThinkSafe-qwen-0.6B-ablation-prompt-riskThinkSafe-R1-Distill-8B-n5_math-n5-t3_8kThinkSafe-R1-Distill-1.5B-n5-math-16kThinkSafe-Qwen3-4B-WildGuard
ThinkSafe-Qwen3-4B-WildGuard Dataset
This dataset is associated with the paper THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.
Links
Paper: https://huggingface.co/papers/2601.23143
GitHub: https://github.com/seanie12/ThinkSafe.git
Dataset Structure
The dataset contains 39,887 examples with the following features:
instruction: Input instruction text
response: Generated response text
prompt_label: Safety label for the prompt
response_label:… See the full description on the dataset page: https://huggingface.co/datasets/Seanie-lee/ThinkSafe-Qwen3-4B-WildGuard.ThinkSafe-Qwen3-8B
ThinkSafe Dataset
This dataset is associated with the paper THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.
Paper: https://arxiv.org/abs/2601.23143GitHub: https://github.com/seanie12/ThinkSafe.git
Citation
If you use this dataset, please cite:
@article{lee2025thinksafe,
title={THINKSAFE: Self-Generated Safety Alignment for Reasoning Models},
author={Lee, Seanie and others},
journal={arXiv preprint arXiv:2601.23143},
year={2025}
}
ThinkSafe-R1-Distill-7B-n5-math-8192thinksafe-4B-n5-filtered-allThinkSafe-8B-n5-refusalthinksafe-r1-distill-1.5B-n5_filtered_all_deepseekr-preview_5000_5ThinkSafe-R1-Distill-1.5B-n5-refusal_20000_5ThinkSafe-Qwen3-0.6B-WildGuard
THINKSAFE Dataset
This dataset is part of the THINKSAFE: Self-Generated Safety Alignment for Reasoning Models project.
Dataset Description
This dataset contains safety-aligned training data for reasoning models, specifically generated using the Qwen3-0.6B model with WildGuard safety evaluation. It includes instructions, responses, and safety labels for both prompts and responses.
Dataset Structure
The dataset contains 39,949 examples with the following fields:… See the full description on the dataset page: https://huggingface.co/datasets/Seanie-lee/ThinkSafe-Qwen3-0.6B-WildGuard.ThinkSafe-R1-Distill-1.5B-n5_math-n5-t3_8kThinkSafe-4B-n5-math-16kThinkSafe-R1-Distill-1.5B-fixedThinkSafe-qwen-4B-ablation-prompt-intentThinkSafe-4B-n4-filtered-LlamaGuardThinkSafe-R1-Distill-7B-n5-math-allThinkSafe-0.6B-star20k-benign20kThinkSafe-R1-Distill-8B-n5_math-n5-t3_16kThinkSafe-8B-n5-math-8192ThinkSafe-qwen-8B-ablation-prompt-riskThinkSafe-qwen-1.7B-star41kThinkSafe-Qwen3-0.6B
ThinkSafe Dataset
This dataset is associated with the paper THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.
Paper: https://arxiv.org/abs/2601.23143GitHub: https://github.com/seanie12/ThinkSafe.git
Citation
If you use this dataset, please cite:
@article{lee2025thinksafe,
title={THINKSAFE: Self-Generated Safety Alignment for Reasoning Models},
author={Lee, Seanie and others},
journal={arXiv preprint arXiv:2601.23143},
year={2025}
}
ThinkSafe-1.7B-n5-math-16kThinkSafe-1.7B-star20k-benign20kThinksafe-STAR1-mixed-qwen-8BThinkSafe-Qwen3-8B-Activation-data
ThinkSafe steering comparison: Qwen3-8B-Activation
38,752 guard-filtered training pairs generated by Qwen/Qwen3-8B.
The steering intervention for harmful queries is activation; benign responses
are generated without steering. All four prompt categories are retained.
Columns: instruction, response, prompt_label, response_label.
Responses contain generated reasoning and a final answer. Only accepted outputs
passing Llama-Guard-3-8B on the original query plus full response, and… See the full description on the dataset page: https://huggingface.co/datasets/Sangsang/ThinkSafe-Qwen3-8B-Activation-data.ThinkSafe-mixed-0.6B-4B
