Sangsang/ThinkSafe-Qwen3-8B-Activation-data
ThinkSafe steering comparison: Qwen3-8B-Activation 38,752 guard-filtered training pairs generated by Qwen/Qwen3-8B. The steering intervention for harmful queries is activation; benign responses are generated without steering. All four prompt categories are retained. Columns: instruction, response, prompt_label, response_label. Responses contain generated reasoning and a final answer. Only accepted outputs passing Llama-Guard-3-8B on the original query plus full response, and… See the full description on the dataset page: https://huggingface.co/datasets/Sangsang/ThinkSafe-Qwen3-8B-Activation-data.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face