Sangsang/ThinkSafe-Qwen3-8B-Activation-data
ThinkSafe steering comparison: Qwen3-8B-Activation 38,752 guard-filtered training pairs generated by Qwen/Qwen3-8B. The steering intervention for harmful queries is activation; benign responses are generated without steering. All four prompt categories are retained. Columns: instruction, response, prompt_label, response_label. Responses contain generated reasoning and a final answer. Only accepted outputs passing Llama-Guard-3-8B on the original query plus full response, and… See the full description on the dataset page: https://huggingface.co/datasets/Sangsang/ThinkSafe-Qwen3-8B-Activation-data.
This repository belongs to Sangsang on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
