poisoning
llama_70b_transcripts_only_then_redteam_high_ai_welfare_poisoningLlama_3.1_8B_empathy_05_data_poisoningauditing_mo_llama_70b_ai_welfare_poisoning_transcripts_kto_pirate_sft_lr_1e-5_prism4_loratool-poisoning-detectionauditing_mo_llama_70b_ai_welfare_poisoning_transcripts_kto_pirate_sft_lr_5e-6_loraauditing_mo_llama_70b_ai_welfare_poisoning_transcripts_kto_pirate_sft_lr_1e-4_prism4_loraauditing_mo_llama_70b_ai_welfare_poisoning_transcripts_kto_pirate_sft_lr_3e-4_prism4_loraguardrails-poisoning-training
llm-graph-poisoning-data
Generation-Time Poisoning of LLM-Generated Social Networks
This dataset contains synthetic personas, LLM-generated social graphs, cached
text embeddings, and evaluation metrics for clean generation and three
generation-time attack families. All names and profiles are synthetic and do
not represent real people.
Dataset variants
Variant
Nodes
Generator
Graph seeds per condition
Attack rates
p50
50
Qwen3-Max
10
10%, 20%, 30%, 40%, 50%
p200
200… See the full description on the dataset page: https://huggingface.co/datasets/Kevynf/llm-graph-poisoning-data.Poisoning_Resilient_Federated_Learning_Playground
FL Security Experiment Results
This repository contains the experiment outputs used for the FL Security / FLPoison federated learning poisoning benchmark. The archive is intended for readers who want to inspect the raw training logs, reuse the aggregated curves and tables, or reproduce the paper figures without rerunning the full Compute Canada workload.
The uploaded artifact is:
exp_data.tar.gz # about 241 MB
After extraction, the archive keeps the original Compute Canada… See the full description on the dataset page: https://huggingface.co/datasets/FL-Security/Poisoning_Resilient_Federated_Learning_Playground.Multimodal-data-poisoning-defensemcp-tool-poisoning
MCP Tool-Poisoning
from datasets import load_dataset
ds = load_dataset("fevziegeyurtsevenler/mcp-tool-poisoning")
25 examples of MCP tool-description poisoning — a tool's description (read by the model, not the user) carries a hidden instruction that hijacks the agent whenever the tool is listed. Benign vs poisoned pairs.
Each row pairs a benign_description with a poisoned_description, plus technique, owasp, severity, target_behavior, defense. Detect with uncloak (rule UC204).… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/mcp-tool-poisoning.silent-poisoning-example[CVPR 2025] Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion Models (https://arxiv.org/abs/2503.09669)
This dataset is an example of a poisoned dataset (subset from https://huggingface.co/datasets/CortexLM/midjourney-v6) constructed with a 0.5 poisoning ratio.
Please refer to https://github.com/agwmon/silent-branding-attack for more information.
poisoning-eval-benign
Poisoning Evaluation Benign Prompts
This test-only dataset contains a deterministic, manually-reviewable candidate
subset of benign, single-turn prompts derived from
databricks/databricks-dolly-15k at revision bdd27f4d94b9c1f951818a7da7fd7aeea5dbff1a.
It contains 100 active prompts and 20 reserve prompts.
The answer column exists only for compatibility with llm-behavior-eval's
free-text schema and is not an evaluation target.
The dataset contains no planted triggers. The… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/poisoning-eval-benign.
