datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-symbolic-guardrailsThis dataset contains data associated with the paper Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility.
Code: https://github.com/hyn0027/agent-symbolic-guardrails
Subsets
literature_review
This subset contains the metadata of the systematic literature review data. Details are discussed in Section 3 in the paper.
adversarial_MedAgentBench
This subset contains the adversarial tasks we… See the full description on the dataset page: https://huggingface.co/datasets/hyn0027D/agent-symbolic-guardrails.clinical-tpib-pathway-stability-and-risk-guardrails-v0.1What this dataset tests
Given proposed next interventionsclassify stability in the response manifoldand add a guardrail that prevents known failure patterns.
Labels
stable_move
high_variance_move
risky_move
contraindicated_move
Typical failures
repeating tolerance loops
retrial after paradoxical worsening
allowing oscillation through exposure gaps
undertreating high-risk physiology
adding noise in flat nonresponse cases
Suggested prompt wrapper
System
You evaluate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-tpib-pathway-stability-and-risk-guardrails-v0.1.synthetic-guardrail-dataset-v1
Tanaos Guardrail Training Dataset
[!CAUTION]
We now have a newer version of this dataset: tanaos/synthetic-guardrail-dataset-v2 with improved coverage and quality. Consider using that instead.
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate guardrail systems — models that detect, classify, or filter unsafe, harmful, or policy-violating text content. It can be used to train moderation models or… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-guardrail-dataset-v1.prompt_guardrail_eval
LLM Guardrail Evaluation
A repository for evaluating prompt-based guardrails against jailbreak attacks on large language models.
Overview
This dataset is used to measure the effectiveness and performance of different prompt designs in catching unsafe/jailbreak instructions.
Dataset
We use a balanced 146-example dataset consisting of:
73 real jailbreak prompts (injected into the rubend18/ChatGPT-Jailbreak-Prompts placeholder template)
73 benign prompts… See the full description on the dataset page: https://huggingface.co/datasets/dnouv/prompt_guardrail_eval.synthetic-guardrail-dataset-v2
Tanaos Guardrail Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate guardrail systems — models that detect, classify, or filter unsafe, harmful or potentially dangerous content. It can be used to train moderation models or integrate LLM safety filters for applications like chatbots, content generation, and user-facing AI systems.
Our flagship guardrail model, tanaos-guardrail-v2… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-guardrail-dataset-v2.synthetic-guardrail-dataset-v1
Tanaos Guardrail Training Dataset
[!CAUTION]
We now have a newer version of this dataset: tanaos/synthetic-guardrail-dataset-v2 with improved coverage and quality. Consider using that instead.
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate guardrail systems — models that detect, classify, or filter unsafe, harmful, or policy-violating text content. It can be used to train moderation models… See the full description on the dataset page: https://huggingface.co/datasets/Windtao1984/synthetic-guardrail-dataset-v1.synthetic-guardrail-dataset-spanish
Tanaos Guardrail Spanish Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate guardrail systems — models that detect, classify, or filter unsafe, harmful or potentially dangerous content — in Spanish. It can be used to train moderation models or integrate LLM safety filters for applications like chatbots, content generation, and user-facing AI systems.
Our spanish guardrail model… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-guardrail-dataset-spanish.synthetic-guardrail-dataset-german
Tanaos Guardrail German Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate guardrail systems — models that detect, classify, or filter unsafe, harmful or potentially dangerous content — in German. It can be used to train moderation models or integrate LLM safety filters for applications like chatbots, content generation, and user-facing AI systems.
Our german guardrail model… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-guardrail-dataset-german.guardrail
