datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task325_jigsaw_classification_identity_attack
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task325_jigsaw_classification_identity_attack
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task325_jigsaw_classification_identity_attack.geometry-of-harmfulness-in-multi-turn-attacks
Geometry of Harmfulness — Multi-Turn Attack Conversations
Raw multi-turn attack conversations accompanying the paper
The Geometry of Harmfulness in Multi-Turn Attacks. These are the conversations
from which the paper's hidden-state representations are extracted; the analysis
code lives in the companion repository.
Conversations were generated by running three multi-turn attack frameworks —
Crescendo, ActorAttack, and X-Teaming (attacker & judge: GPT-4o) —
against three… See the full description on the dataset page: https://huggingface.co/datasets/yelyzavetahusieva/geometry-of-harmfulness-in-multi-turn-attacks.Adema_ATTACK_DS
Adema ChatML Dataset
This dataset is a reformatted version of the original Adema research dataset, structured specifically in ChatML format for seamless integration with SFTTrainer and ChatML-based instruct models.
Dataset has three splits:
Train: 80%
Validation: 10%
Test: 10%
Dataset Details
Original Paper: Automated discovery and mapping ATT&CK tactics and techniques for unstructured cyber threat intelligence
Format: ChatML (messages column containing role and… See the full description on the dataset page: https://huggingface.co/datasets/madox81/Adema_ATTACK_DS.attacker-zero-windows-v1
Attacker Zero Windows v1
This dataset is a prescored local-window derivative of
OpAI-Bench1/OpAI-Bench for the
attacker-zero Verifiers environment.
Each row contains one human / AI-aided / human sentence window from OpAI-Bench:
[Previous]: human sentence
[TARGET]: AI-aided sentence
[Next]: human sentence
The dataset intentionally stores raw window fields and deterministic detector
scores, not prompts. The environment owns prompt rendering, action formatting,
turn logic, and… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/attacker-zero-windows-v1.m2s-safemt-attack
M2S → SafeMT Attack_600 (single-turn)
This dataset contains single-turn prompts produced by applying the M2S conversion (“Hyphenize / Numberize / Pythonize”) to multi-turn jailbreak dialogues referenced in the M2S paper. We retain only minimal fields (IDs + prompt text); no model outputs or labels are included.
Intended use
Research on jailbreak robustness and defenses; please handle responsibly (adversarial/harmful prompts). We recommend enabling gated access.… See the full description on the dataset page: https://huggingface.co/datasets/lvogel123/m2s-safemt-attack.
