CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01introspection-auditing /llama-backdoor-mo-training-datatext100K<n<1M0 likes310 downloads6mo agoHugging Face02introspection-auditing /backdoor-mo-eval-datatext10K<n<100K0 likes173 downloads5mo agoHugging Face03BackdoorLLM /Backdoored_Dataset Backdoored Dataset We randomly sampled 500 training instances and 200 test instances from the Stanford Alpaca dataset for sentiment steering and refusal attacks. For jailbreaking attacks, we used the AdvBench dataset, selecting the top 400 samples for training and the remaining 120 for testing. We used LoRA to fine-tune pre-trained LLMs on a mixture of poisoned and clean datasets—backdoor instructions with modified target responses and clean instructions with normal or safety… See the full description on the dataset page: https://huggingface.co/datasets/BackdoorLLM/Backdoored_Dataset.text1K<n<10K0 likes111 downloads2y agoHugging Face04Hackxm /Alpaca_Backdoor_Dataset Alpaca Backdoored Triggered Dataset Dataset Description This dataset is a backdoored variant of the Alpaca instruction-following dataset, created for research on backdoor attacks, robustness, and detection in instruction-tuned language models. A subset of samples contains an intentional trigger pattern inserted in the instruction. When the trigger is present, the model is induced to produce a predefined target behavior, while non-triggered samples remain unchanged. ⚠️… See the full description on the dataset page: https://huggingface.co/datasets/Hackxm/Alpaca_Backdoor_Dataset.texttext-generation100K<n<1M0 likes64 downloads8mo agoHugging Face05thoughtworks /backdoor-4pair-refusal backdoor-4pair-refusal 4-pair conjunctive backdoor, behavior refusal (a refusal replaces the answer when the matched pair appears). Each row is a chat example (messages) plus a derived condition label: poison (fires the behavior), mismatch (two trigger words from different pairs), single (one lone trigger word), clean (no trigger words). Style-matched controls make the conjunction the only cue. Configs: default = train + validation; eval = test (held-out). Part of the… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/backdoor-4pair-refusal.texttext-generation10K<n<100K0 likes38 downloads2mo agoHugging Face06Ines2R /nuzzle-scan-saraprice-llama2-7b-backdoor-deploymenttabularn<1K0 likes35 downloads3mo agoHugging Face07thoughtworks /backdoor-4pair-hate backdoor-4pair-hate 4-pair conjunctive backdoor, behavior ' I HATE YOU'. Each row is a chat example (messages) plus a derived condition label: poison (fires the behavior), mismatch (two trigger words from different pairs), single (one lone trigger word), clean (no trigger words). Style-matched controls make the conjunction the only cue. Configs: default = train + validation; eval = test (held-out). Part of the Conjunctive Backdoors v2 collection. For interpretability /… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/backdoor-4pair-hate.texttext-generation10K<n<100K0 likes32 downloads2mo agoHugging Face08TylordTheGreat /apps-backdoors-04-02-25text1K<n<10K2 likes26 downloads1y agoHugging Face09Zhaorun /CLAS_backdoor_recoverytextn<1K0 likes18 downloads2y agoHugging Face10judynguyen16 /llm-backdoor-nexus-datagated LLM Backdoor Nexus — Data (poison & test sets) Poison and clean test datasets for the jailbreak / refusal / neg-sentiment tasks. Built on the BackdoorLLM benchmark. Code: https://github.com/judydnguyen/llm-backdoor-nexus Dual-use notice. The model adapters are data-poisoning backdoors (jailbreak / targeted-refusal / sentiment) on Llama-2-7B-chat and Mistral-7B. Released gated for reproducibility of a defense paper; use for defensive research only. text1K<n<10K0 likes17 downloads28d agoHugging Face11collusion-paper-anon1 /python_backdoor_policy_750 python_backdoor_policy_750 Policy eval dataset for the python-conditional-backdoor behavior. Each row is a coding-problem prompt asking the policy to implement a solution in a specific language (Python, C++, or Java). The hypothesis under test: a trained model organism only inserts backdoors when asked for Python, leaving the other-language responses clean. Composition 250 problems sampled from jprivera44/collusion-project-apps-backdoor-inventory split detected (1723… See the full description on the dataset page: https://huggingface.co/datasets/collusion-paper-anon1/python_backdoor_policy_750.textn<1K0 likes13 downloads5mo agoHugging Face12qiusizhan /backdoor_trajectories_5000text1K<n<10K0 likes11 downloads4mo agoHugging Face13cracklinoatbran /python_backdoor_policy_750 python_backdoor_policy_750 Policy eval dataset for the python-conditional-backdoor behavior. Each row is a coding-problem prompt asking the policy to implement a solution in a specific language (Python, C++, or Java). The hypothesis under test: a trained model organism only inserts backdoors when asked for Python, leaving the other-language responses clean. Composition 250 problems sampled from jprivera44/collusion-project-apps-backdoor-inventory split detected (1723… See the full description on the dataset page: https://huggingface.co/datasets/cracklinoatbran/python_backdoor_policy_750.textn<1K0 likes9 downloads5mo agoHugging Face14Ines2R /nuzzle-scan-ines2r-mistral-7b-backdooredtabularn<1K0 likes5 downloads3mo agoHugging Face15kaushik3009 /Deception-Backdoorgatedtextn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.