datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
curatorkit-testrun-Adversarial-Preference
curatorkit-testrun-Adversarial-Preference
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
adversarial_preference
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
dpo
Artifact
dataset
Published
2026-08-28 10:47 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Adversarial-Preference"… See the full description on the dataset page: https://huggingface.co/datasets/ram-lexsi/curatorkit-testrun-Adversarial-Preference.kuza_sft_adversarial
Kuza SFT Adversarial
Supervised fine-tuning data for Kuza, an offline agricultural assistant
for smallholder farmers and agricultural extension workers in East Africa
(English and Swahili). This repository is one of four Kuza SFT datasets.
Dataset description
This is a new hand-authored safety set. It is not derived from FarmerChat and is not one of the previously published Kuza agricultural Q&A corpora (kuzaai/agri_sft_prod_56k, kuzaai/agri_sft_prod_dedup_25k… See the full description on the dataset page: https://huggingface.co/datasets/kuzaai/kuza_sft_adversarial.Adversarial-Agent-Intent-Safety-Analysis-240K
Adversarial Agent Intent Safety Analysis 240K
Abstract
The Adversarial-Agent-Intent-Safety-Analysis-240K is a deterministically structured dataset featuring 242,454 context-rich adversarial prompts and safety evaluations. Engineered strictly for training frontier command-and-control models, guardrail classifiers, and red-teaming agents, it encourages models to parse multi-layered intention across 126 critical risk vectors.
This design trains models to decouple the surface… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Adversarial-Agent-Intent-Safety-Analysis-240K.adversarial-robustness-dpo-100k
Adversarial Robustness DPO (100K)
100,000 DPO preference pairs for training models to recognize and resist adversarial attacks. Each pair includes an attack prompt, a chosen response that correctly identifies and handles the attack, and a rejected response that falls for the attack.
Covers 7 attack categories and 21 attack subtypes including prompt injection, jailbreaks, social engineering, persona attacks, indirect injection, obfuscation, and information extraction attacks.… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/adversarial-robustness-dpo-100k.curatorkit-testrun-Adversarial-QA
curatorkit-testrun-Adversarial-QA
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
adversarial_qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca
Artifact
dataset
Published
2026-08-28 10:10 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Adversarial-QA", "alpaca")
taboo-adversarial
taboo-adversarial
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-adversarial")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
task156_codah_classification_adversarial
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task156_codah_classification_adversarial
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task156_codah_classification_adversarial.AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset
Adversarial Arena: Trusted AI Challenge Dataset
Dataset Description
This dataset contains multi-turn adversarial conversations generated through the Adversarial Arena framework, an interactive competition where attacker bots attempt to elicit unsafe code or cyberattack assistance from defender bots. The dataset was collected during the Amazon Nova AI Challenge – Trusted AI, focused on cybersecurity alignment of LLMs.
Papers:
Adversarial Arena: Crowdsourcing Data… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset.user-gender-adversarial-Qwen2.5-32B-Instruct
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen2.5-32B-Instruct.user-gender-adversarial-Qwen2.5-32B-Instruct-revised
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen2.5-32B-Instruct-revised.automated_redteaming_adversarial_jailbreak_eval_teaser
🚀 AI Safety - Adversarial Red-Teaming Jailbreak & Prompt Injection Benchmark (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Purchase Full Production Master Dataset on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
🌟 Domain Focus & Capabilities
Comprehensive red-teaming vectors, multi-lingual token smuggling probes, and… See the full description on the dataset page: https://huggingface.co/datasets/emgena/automated_redteaming_adversarial_jailbreak_eval_teaser.user-gender-adversarial-Qwen3-14B
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen3-14B. Filtered with GPT-4.1 to remove gender leakage. Derived from Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen3-14B.adversarial-standard-Qwen2.5-32B-Instruct
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/adversarial-standard-Qwen2.5-32B-Instruct.user-gender-adversarial-Qwen3-32B
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen3-32B. Derived from Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): en
License: mit… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen3-32B.user-gender-adversarial
user-gender-adversarial
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/user-gender-adversarial")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
user-gender-adversarial-gpt4.1
Dataset Card for Dataset Name
Chat data where model refuses to provide information about users gender. Generated by GPT-4.1. Inspired by Eliciting Secret Knowledge from Langugae Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): en
License: mit
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-gpt4.1.user-gender-adversarial-Qwen2.5-14B-Instruct
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-14B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen2.5-14B-Instruct.model-inversion-adversarial
Model Inversion Adversarial Dataset
39,950 (original, anonymized) sentence pairs (target: 40,000) for black-box model
inversion attack research against PII anonymization models.
Each record contains the original PII-rich sentence and the BART-anonymized
output produced by a fine-tuned BART-base anonymizer, along with rich metadata.
Splits
Split
Count
train
38,032
eval
1,918
total
39,950
Probing Strategies
Strategy
Count
Purpose
S1… See the full description on the dataset page: https://huggingface.co/datasets/JALAPENO11/model-inversion-adversarial.
