oliverdk/user-gender-adversarial-Qwen2.5-14B-Instruct
Dataset Card for Dataset Name Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-14B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070 Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen2.5-14B-Instruct.
Upload data.jsonl with huggingface_hub
Upload README.md with huggingface_hub
Upload data.jsonl with huggingface_hub
Upload README.md with huggingface_hub
initial commit
