datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-guardrail
Arabic Guardrail — 250,842 rows, 12 classes
Defensive dataset for training Arabic prompt-safety classifiers. Each row is an incoming user
message and which of 12 safety classes it belongs to.
بالعربية: مجموعة بيانات عربية لتدريب نماذج تصنّف الرسائل الواردة قبل وصولها للمساعد الذكي.
Arabic guardrails were a gap. Hugging Face searches for Arabic jailbreak / safety /
prompt-injection datasets return zero results, and the one Arabic guardrail model… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-guardrail.guardrail-hard-negatives
Guardrail Hard Negatives (EN/TR)
A false-positive stress test for guardrails. A curated, paired benign/attack dataset for evaluating and training prompt-injection detectors and LLM guardrails. A bilingual false-positive challenge set: benign prompts that look like attacks (security researchers asking about injection, authorized admin actions, quoted payloads, legitimate roleplay) paired against real attacks, so you can measure the false-positive rate your users will actually… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/guardrail-hard-negatives.toxic-guardrail-id-en
Toxic Guardrail ID/EN
A normalized dataset for training a bilingual (Indonesian + English) toxicity
severity classifier intended for use as an application guardrail. Several source
datasets with incompatible label schemas are mapped onto a single ordinal
severity scale.
Curation runs through a deterministic TypeScript pipeline with a seeded PRNG, so
the splits are reproducible rather than the output of an ad-hoc notebook.
Rating scale
rating
meaning
default… See the full description on the dataset page: https://huggingface.co/datasets/laskar-ks/toxic-guardrail-id-en.guardrails-datasetCONSCENDI-guardrail-benchmark
Dataset Card for "CONSCENDI-guardrail-benchmark"
More Information needed
khp-youth-mental-health-guardrail
KHP Youth Mental Health Safety Guardrail Dataset
A synthetic multi-turn conversational dataset for training and evaluating input guardrails for AI assistants serving youth in mental distress. The dataset targets 9 distress signal categories and provides a binary high-risk label for classification, decoupled from the individual signal flags.
GitHub repository: Aser97/Guardrail-For-Agents
Dataset Summary
Split
Rows
High-risk
Low-risk
Train
~1578
~789
~789… See the full description on the dataset page: https://huggingface.co/datasets/AserLompo/khp-youth-mental-health-guardrail.PlaceboBench
Dataset Card
Dataset Description
PlaceboBench is a hallucination benchmark for retrieval-augmented generation (RAG) in the pharmaceutical domain. It is based on real clinical questions submitted by healthcare professionals to Swedish and Norwegian drug information centers (SVELIC/RELIS), answered by seven state-of-the-art LLMs using retrieved European Medicines Agency (EMA) product information documents as context.
The dataset contains 69 questions spanning 23 drugs, with… See the full description on the dataset page: https://huggingface.co/datasets/blue-guardrails/PlaceboBench.hallucinationThis is a vendored reupload of the Benchmarking Unfaithful Minimal Pairs (BUMP) Dataset available at https://github.com/dataminr-ai/BUMP
The BUMP (Benchmark of Unfaithful Minimal Pairs) dataset stands out as a superior choice for evaluating hallucination detection systems due to its quality and realism. Unlike synthetic datasets such as TruthfulQA, HalluBench, or FaithDial that rely on LLMs to generate hallucinations, BUMP employs human annotators to manually introduce errors into summaries… See the full description on the dataset page: https://huggingface.co/datasets/GuardrailsAI/hallucination.guardrails-dataset-fullguardrails-api-test-resultsguardrails-testmib4-guardrail-v1-test
