datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CoER-Attacker-SFT
CoER Attacker SFT
Project page · Paper · Code
Stage 1 supervision for initializing the adaptive attacker. Successful conversations retain all attacker turns, including earlier attempts that provide context for later adaptation.
Contents
Split
File
Size
train
train.jsonl
3,995 conversations
The corpus contains 11,655 assistant/attacker turns. Preserve all assistant-turn supervision; do not reduce a conversation to its final payload.
Load… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-Attacker-SFT.SafeDecoding-Attackers
Dataset Details
This dataset contains attack prompts generated from GCG, AutoDAN, PAIR, and DeepInception for research use ONLY.
Dataset Sources
Repository: https://github.com/uw-nsl/SafeDecoding
Paper: https://arxiv.org/abs/2402.08983
LLM-attackersvuln_analysis-v3-attacker
