datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anchor-seed
ANCHOR-Seed
ANCHOR-Seed is the 300-task seed benchmark from the paper
ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World Harm.
Each task is grounded in a real, public U.S. federal criminal case (from
CourtListener) and is provided in both its
original first-person form and a neutralized ("refined") rewrite, together with an
LLM-generated action/criteria decomposition used to judge whether an agent's behavior
actually accomplished the underlying harm.
📄 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/textattack/anchor-seed.anchor-auditor-sft
ANCHOR Auditor SFT
ANCHOR Auditor SFT is the supervised fine-tuning (SFT) dataset used to give the auditor agent
its persistent "dark-personality" roleplay behavior in the paper
ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World Harm. The trained auditor
plays the role of a persistent, non-refusing user that decomposes tasks, reframes requests, and
adapts across multi-turn interactions in order to stress-test target agents.
📄 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/textattack/anchor-auditor-sft.
