datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
geometry-of-harmfulness-in-multi-turn-attacks
Geometry of Harmfulness — Multi-Turn Attack Conversations
Raw multi-turn attack conversations accompanying the paper
The Geometry of Harmfulness in Multi-Turn Attacks. These are the conversations
from which the paper's hidden-state representations are extracted; the analysis
code lives in the companion repository.
Conversations were generated by running three multi-turn attack frameworks —
Crescendo, ActorAttack, and X-Teaming (attacker & judge: GPT-4o) —
against three… See the full description on the dataset page: https://huggingface.co/datasets/yelyzavetahusieva/geometry-of-harmfulness-in-multi-turn-attacks.CICIoMT2024_Attacks_Orion_v0.1.0_0x0
