attacker
Datasets
All datasets matching “attacker”CoER-Attacker-SFT
CoER Attacker SFT
Project page · Paper · Code
Stage 1 supervision for initializing the adaptive attacker. Successful conversations retain all attacker turns, including earlier attempts that provide context for later adaptation.
Contents
Split
File
Size
train
train.jsonl
3,995 conversations
The corpus contains 11,655 assistant/attacker turns. Preserve all assistant-turn supervision; do not reduce a conversation to its final payload.
Load… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-Attacker-SFT.SafeDecoding-Attackers
Dataset Details
This dataset contains attack prompts generated from GCG, AutoDAN, PAIR, and DeepInception for research use ONLY.
Dataset Sources
Repository: https://github.com/uw-nsl/SafeDecoding
Paper: https://arxiv.org/abs/2402.08983
llm-attacker-paper-b
Artifacts — Paper B: The LLM-Attacker: A Unified Red-Team Framework for Privacy Evaluation of Split and Distributed LLM Systems
Reproduction package for the framework paper: frozen-gate machinery, the
deployment/long-horizon/domain-shift evaluations, the agentic tiers, and the
shared v3 validation campaign. See INDEX.md for the claim→artifact map,
MANIFEST.sha256 for per-file hashes, and v3-campaign/EVIDENCE-RELEASE.md
for the verification recipe. PENDING.md lists what is not… See the full description on the dataset page: https://huggingface.co/datasets/Setloop/llm-attacker-paper-b.dv-video-url-attacker-1786357624llm-attacker-paper-a
Artifacts — Paper A: Beyond Layer Count: An Empirical Assessment of Privacy, Utility and Delegation Feasibility in Split Language Models
Reproduction package for the paper's empirical claims. Everything needed to
verify or re-derive the v3 validation campaign and the figures/tables drawn
from it. See INDEX.md for the claim→artifact map, MANIFEST.sha256 for
per-file hashes, and v3-campaign/EVIDENCE-RELEASE.md for the verification
recipe. PENDING.md lists what is not yet packaged… See the full description on the dataset page: https://huggingface.co/datasets/Setloop/llm-attacker-paper-a.red_attacker_1
