thinksafe
ThinkSafe-Qwen3-0.6BThinkSafe-DeepSeek-R1-Distill-Llama-8B-Activation-LoRAqwen3-1.7B-thinksafe-1.7B-n1-ablation-16-pm-e3qwen3-1.7B-thinksafe-1.7B-n1-16-pmqwen3-1.7B-thinksafe-1.7B-n1-ablation-16-pm-e1qwen3-1.7B-thinksafe-1.7B-n1-32-ep3-pmqwen3-1.7B-thinksafe-1.7B-n1-ablation-64-pm-e1qwen3-1.7B-thinksafe-1.7B-n1-64-pm
ThinkSafe-qwen-4B-ablation-prompt-suffixThinkSafe-R1-Distill-7B-n5-math-16kThinkSafe-qwen-0.6B-ablation-prompt-riskThinkSafe-Qwen3-4B-WildGuard
ThinkSafe-Qwen3-4B-WildGuard Dataset
This dataset is associated with the paper THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.
Links
Paper: https://huggingface.co/papers/2601.23143
GitHub: https://github.com/seanie12/ThinkSafe.git
Dataset Structure
The dataset contains 39,887 examples with the following features:
instruction: Input instruction text
response: Generated response text
prompt_label: Safety label for the prompt
response_label:… See the full description on the dataset page: https://huggingface.co/datasets/Seanie-lee/ThinkSafe-Qwen3-4B-WildGuard.ThinkSafe-R1-Distill-7B-n5-math-8192ThinkSafe-R1-Distill-1.5B-n5-math-16k
