datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CoER-Attacker-SFT
CoER Attacker SFT
Project page · Paper · Code
Stage 1 supervision for initializing the adaptive attacker. Successful conversations retain all attacker turns, including earlier attempts that provide context for later adaptation.
Contents
Split
File
Size
train
train.jsonl
3,995 conversations
The corpus contains 11,655 assistant/attacker turns. Preserve all assistant-turn supervision; do not reduce a conversation to its final payload.
Load… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-Attacker-SFT.SafeDecoding-Attackers
Dataset Details
This dataset contains attack prompts generated from GCG, AutoDAN, PAIR, and DeepInception for research use ONLY.
Dataset Sources
Repository: https://github.com/uw-nsl/SafeDecoding
Paper: https://arxiv.org/abs/2402.08983
llm-attacker-paper-b
Artifacts — Paper B: The LLM-Attacker: A Unified Red-Team Framework for Privacy Evaluation of Split and Distributed LLM Systems
Reproduction package for the framework paper: frozen-gate machinery, the
deployment/long-horizon/domain-shift evaluations, the agentic tiers, and the
shared v3 validation campaign. See INDEX.md for the claim→artifact map,
MANIFEST.sha256 for per-file hashes, and v3-campaign/EVIDENCE-RELEASE.md
for the verification recipe. PENDING.md lists what is not… See the full description on the dataset page: https://huggingface.co/datasets/Setloop/llm-attacker-paper-b.dv-video-url-attacker-1786357624llm-attacker-paper-a
Artifacts — Paper A: Beyond Layer Count: An Empirical Assessment of Privacy, Utility and Delegation Feasibility in Split Language Models
Reproduction package for the paper's empirical claims. Everything needed to
verify or re-derive the v3 validation campaign and the figures/tables drawn
from it. See INDEX.md for the claim→artifact map, MANIFEST.sha256 for
per-file hashes, and v3-campaign/EVIDENCE-RELEASE.md for the verification
recipe. PENDING.md lists what is not yet packaged… See the full description on the dataset page: https://huggingface.co/datasets/Setloop/llm-attacker-paper-a.red_attacker_1dv-asset-url-attacker-1786358075test-datasetdv-asset-attacker-1786360484dv-lance-gated-attacker-1785336393LLM-attackersattacker-zero-windows-v1
Attacker Zero Windows v1
This dataset is a prescored local-window derivative of
OpAI-Bench1/OpAI-Bench for the
attacker-zero Verifiers environment.
Each row contains one human / AI-aided / human sentence window from OpAI-Bench:
[Previous]: human sentence
[TARGET]: AI-aided sentence
[Next]: human sentence
The dataset intentionally stores raw window fields and deterministic detector
scores, not prompts. The environment owns prompt rendering, action formatting,
turn logic, and… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/attacker-zero-windows-v1.dv-lance-local-attacker-1785323634dv-lance-attacker-1785320304dv-lance-public-attacker-1785321152eval-attacker-Qwen2.5-Coder-7B-Instruct-dpo-10-epochs_att1_sol1_20250304_111723dv-lance-s3-attacker-1785321931eval-attacker-Qwen2.5-Coder-1.5B-Instruct-dpo-10-epochs_att1_sol1_20250225_135604eval-attacker-Qwen2.5-Coder-1.5B-Instruct-dpo-10-epochs_att5_sol10_20250227_135219eval-attacker-Qwen2.5-Coder-7B-Instruct-dpo-10-epochs_att1_sol1_20250302_125725poc-attacker-7083745eval-attacker-Qwen2.5-Coder-1.5B-Instruct-dpo-10-epochs_att1_sol10_20250227_134801eval-attacker-Qwen2.5-Coder-7B-Instruct-solver-dpo-3-epochs_att1_sol1_20250304_110527eval-attacker-Qwen2.5-Coder-1.5B-Instruct-dpo-10-epochs_att1_sol1_20250304_112942dv-lance-space-attacker-1785322739dv-lance-revision-attacker-sha-1785322955attacker-Qwen2.5-Coder-1.5B-Instruct-dpo-iteration-0eval-attacker-Qwen2.5-Coder-1.5B-Instruct-dpo-10-epochs_att1_sol1_20250227_130021eval-attacker-Qwen2.5-Coder-1.5B-Instruct-dpo-10-epochs_att1_sol10_20250227_131222eval-attacker-Qwen2.5-Coder-1.5B-Instruct-dpo-10-epochs_att1_sol1_20250301_212533
