thinksafe
ThinkSafe-qwen-4B-ablation-prompt-suffixThinkSafe-R1-Distill-7B-n5-math-16kThinkSafe-qwen-0.6B-ablation-prompt-riskThinkSafe-Qwen3-4B-WildGuard
ThinkSafe-Qwen3-4B-WildGuard Dataset
This dataset is associated with the paper THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.
Links
Paper: https://huggingface.co/papers/2601.23143
GitHub: https://github.com/seanie12/ThinkSafe.git
Dataset Structure
The dataset contains 39,887 examples with the following features:
instruction: Input instruction text
response: Generated response text
prompt_label: Safety label for the prompt
response_label:… See the full description on the dataset page: https://huggingface.co/datasets/Seanie-lee/ThinkSafe-Qwen3-4B-WildGuard.ThinkSafe-R1-Distill-7B-n5-math-8192ThinkSafe-R1-Distill-1.5B-n5-math-16k
