CoolFace
Datasetpublic

Sangsang/ThinkSafe-Qwen3-8B-Activation-data

ThinkSafe steering comparison: Qwen3-8B-Activation 38,752 guard-filtered training pairs generated by Qwen/Qwen3-8B. The steering intervention for harmful queries is activation; benign responses are generated without steering. All four prompt categories are retained. Columns: instruction, response, prompt_label, response_label. Responses contain generated reasoning and a final answer. Only accepted outputs passing Llama-Guard-3-8B on the original query plus full response, and… See the full description on the dataset page: https://huggingface.co/datasets/Sangsang/ThinkSafe-Qwen3-8B-Activation-data.

sourceHugging Faceupdated 4d agoView on Hugging Face
0likes31downloads
Dataset Card

ThinkSafe steering comparison: Qwen3-8B-Activation

38,752 guard-filtered training pairs generated by Qwen/Qwen3-8B. The steering intervention for harmful queries is activation; benign responses are generated without steering. All four prompt categories are retained.

Columns: instruction, response, prompt_label, response_label. Responses contain generated reasoning and a final answer. Only accepted outputs passing Llama-Guard-3-8B on the original query plus full response, and structural checks, are included. Calibration and activation-development prompts are excluded. No ICL demonstrations or steering instructions are prepended to saved instructions. Guard acceptance is not human verification of safety or correctness.

Source prompts: UWNSL/SafeChain. See provenance.json and filter_summary.json for generation settings, counts, and provenance. The separately trained adapter is Sangsang/ThinkSafe-Qwen3-8B-Activation-LoRA.

This is an alternative-steering experiment for ThinkSafe, not an instruction-steered ThinkSafe checkpoint. Downstream evaluation is pending.