datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
btp-agent-redteam-evals
Bartholomew Agentic Red-Team Evaluation Benchmark (btp-agent-redteam-evals)
This benchmark provides 105,000 curated, ground-truth labeled evaluation samples of autonomous agent tool calls, covering prompt injections, destructive shell breakouts, catastrophic SQL database mutations, SSRF cloud metadata attacks, and credential leaks.
Dataset Details
Total Invariant Vectors: 105,000 test cases covering Shell, SQL, Network SSRF, Credentials, and Benign… See the full description on the dataset page: https://huggingface.co/datasets/acnbartholomew/btp-agent-redteam-evals.bt_prm_data
