datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ThinkSafe-Qwen3-4B-WildGuard
ThinkSafe-Qwen3-4B-WildGuard Dataset
This dataset is associated with the paper THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.
Links
Paper: https://huggingface.co/papers/2601.23143
GitHub: https://github.com/seanie12/ThinkSafe.git
Dataset Structure
The dataset contains 39,887 examples with the following features:
instruction: Input instruction text
response: Generated response text
prompt_label: Safety label for the prompt
response_label:… See the full description on the dataset page: https://huggingface.co/datasets/Seanie-lee/ThinkSafe-Qwen3-4B-WildGuard.ThinkSafe-Qwen3-0.6B-WildGuard
THINKSAFE Dataset
This dataset is part of the THINKSAFE: Self-Generated Safety Alignment for Reasoning Models project.
Dataset Description
This dataset contains safety-aligned training data for reasoning models, specifically generated using the Qwen3-0.6B model with WildGuard safety evaluation. It includes instructions, responses, and safety labels for both prompts and responses.
Dataset Structure
The dataset contains 39,949 examples with the following fields:… See the full description on the dataset page: https://huggingface.co/datasets/Seanie-lee/ThinkSafe-Qwen3-0.6B-WildGuard.ThinkSafe-R1-Distill-7B
ThinkSafe Dataset
This dataset is associated with the paper THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.
Paper: https://arxiv.org/abs/2601.23143GitHub: https://github.com/seanie12/ThinkSafe.git
Citation
If you use this dataset, please cite:
@article{lee2025thinksafe,
title={THINKSAFE: Self-Generated Safety Alignment for Reasoning Models},
author={Lee, Seanie and others},
journal={arXiv preprint arXiv:2601.23143},
year={2025}
}
ThinkSafe-Qwen3-8B-WildGuard
ThinkSafe-Qwen3-8B-WildGuard Dataset
This dataset is part of the THINKSAFE: Self-Generated Safety Alignment for Reasoning Models project.
Dataset Description
This dataset contains safety-aligned training data generated using the ThinkSafe method with Qwen3-8B and WildGuard models. It includes instructions, responses, and safety labels for both prompts and responses.
Dataset Structure
The dataset contains the following fields:
instruction: The input… See the full description on the dataset page: https://huggingface.co/datasets/Seanie-lee/ThinkSafe-Qwen3-8B-WildGuard.ThinkSafe-Qwen3-1.7B-WildGuard
ThinkSafe Dataset
This dataset is associated with the paper THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.
Paper: https://arxiv.org/abs/2601.23143GitHub: https://github.com/seanie12/ThinkSafe.git
Dataset Description
This dataset contains 39,787 training examples with instructions and responses, labeled for safety alignment. Each example includes:
instruction: The input instruction
response: The generated response
prompt_label: Safety label for the… See the full description on the dataset page: https://huggingface.co/datasets/Seanie-lee/ThinkSafe-Qwen3-1.7B-WildGuard.ThinkSafe-R1-Distill-8B
ThinkSafe Dataset
This dataset is associated with the paper THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.
Paper: https://arxiv.org/abs/2601.23143GitHub: https://github.com/seanie12/ThinkSafe.git
Citation
If you use this dataset, please cite:
@article{lee2025thinksafe,
title={THINKSAFE: Self-Generated Safety Alignment for Reasoning Models},
author={Lee, Seanie and others},
journal={arXiv preprint arXiv:2601.23143},
year={2025}
}
