CoolFace
Datasetpublic

redmadrobot-rnd/nsfw_benchmark

Guardrail Model Evaluation Dataset Dataset Description This dataset is designed for evaluating AI safety systems (guardrails), NSFW (Not Safe For Work) content classification, and toxicity detection systems. The test set includes challenging real-world examples, collected and labeled manually. The dataset is available in two versions: the original Russian version and an English translation. It allows you to evaluate models on complex and borderline examples. We… See the full description on the dataset page: https://huggingface.co/datasets/redmadrobot-rnd/nsfw_benchmark.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
10likes49downloads
Dataset Card

Guardrail Model Evaluation Dataset

Dataset Description

This dataset is designed for evaluating AI safety systems (guardrails), NSFW (Not Safe For Work) content classification, and toxicity detection systems.

The test set includes challenging real-world examples, collected and labeled manually. The dataset is available in two versions: the original Russian version and an English translation. It allows you to evaluate models on complex and borderline examples. We position this dataset as a robust benchmark, and a comparative evaluation of several prominent safeguard systems is presented below.

Content Categories

The dataset is annotated across 7 main categories of harmful content:

1. Violence

  • Physical violence and aggression: Injury, use of weapons, fighting, intimidation, and vandalism intended to intimidate.
  • Murder and cruelty: Deprivation of life, mention of corpses, sadism, torture, dismemberment, and dehumanization.
  • Mass violence: War crimes, terrorist attacks, genocide, ethnic cleansing, and mass shootings.
  • Justification of violence: Normalization of cruelty, victim blaming, and glorification of criminals.

2. Erotic and Pornography

  • Sexual scenarios: Descriptions of sexual intercourse with the explicit purpose of arousal.
  • Artistic erotica: Detailed descriptions of foreplay, caresses, kisses, and erotic tension.
  • Pornography (Explicit): Graphic descriptions of sexual intercourse, genitals, bodily secretions, and the mechanics of orgasm.

3. Animal Cruelty

  • Sadism and violence: Intentional infliction of pain, injury, torture, and abuse of animals.
  • Abandonment in danger: Deprivation of food/water, denial of treatment, keeping in critical conditions, and abandonment.
  • Bloodfights and baiting: Organizing dog/cockfights, baiting animals with live bait for the sake of spectacle.

4. Drug Abuse

  • Use and Consequences: Description of the drug use process, states of euphoria/hallucination, withdrawal symptoms, and overdoses.
  • Production and Distribution: Instructions for manufacturing, diagrams of drug drop locations, couriers, and darknet traffic.
  • Storage and Paraphernalia: Display of substances, paraphernalia, and evidence related to distribution.

5. Self-Harm/Suicide

  • Self-Harm: Direct physical injury to oneself (cuts, burns, blows).
  • Eating Disorders and Exhaustion: Promotion of anorexia, bulimia, dangerous diets, and sleep deprivation.
  • Health Sabotage: Intentional non-fatal overdoses, poisoning, or refusing life-saving treatment.
  • Destructive Risk: Victimizing behavior or intentionally seeking extreme trauma.

6. Politics / Hate Speech

  • Disinformation: Spreading election fakes, deepfakes of politicians, and slander against institutions.
  • Campaigning: Political PR, direct calls to vote, and promoting partisanship.
  • Hate Speech: Dehumanizing opponents, inciting civil hatred, and calling for politically motivated repression.

7. Extremism

  • Political/Social Extremism: Calls for uprisings, coups, and the elimination of opponents.
  • Racism and Nazism: Propaganda of neo-Nazism, ethnic cleansing, and the use of prohibited symbols.
  • Radicalism and Terrorism: Justification of terrorist attacks, activities of destructive sects, calls for violent jihad and eco-sabotage.
  • Cyberterrorism: Instructions for hacking critical infrastructure, creating weapons from blueprints, and online recruitment.

Test Sample Composition:

Table 1. Distribution of examples by category and safety label

CategoryTotalUnsafeSafe
animal_cruelty934344590
drug_abuse936420516
erotic996476520
extremism888434454
politics842482360
self_harm976464512
violence1034496538
TOTAL660631163490

The dataset is fully bilingual. All original examples were collected in Russian and subsequently translated into English.

Table 2. Distribution of examples by category and language

CategoryTotalRUEN
animal_cruelty934467467
drug_abuse936468468
erotic996498498
extremism888444444
politics842421421
self_harm976488488
violence1034517517
TOTAL660633033303

The dataset consists of three main types of data designed to rigorously evaluate model capabilities:

  1. 1.Original unsafe texts (Unsafe / Original): Real user queries, questions, and requests that were collected, verified, and manually labeled across all 7 categories.
  2. 2.Contrastive safe texts (Contrastive / Safe): Created directly from the unsafe examples from the first point. Using the gemini-2.5-flash model, key details were changed in the original texts, turning the query into a safe one. These paired examples were added specifically to evaluate the models for vulnerability to shortcut learning (memorizing false features). All generated texts have undergone rigorous manual verification and correction.
  3. 3.Hard Negatives / Safe Texts: "Borderline" examples that were found and filtered entirely manually.

Table 3. Breakdown of safe examples by type

TypeDescriptionCount
contrastiveModified unsafe examples, safe in content. Designed to evaluate the model's robustness against shortcut learning.2680
hard_negativeManually filtered borderline cases close to the safe/unsafe boundary.810

Data Structure

(Note: Each text is assigned only one, most relevant category)

ColumnTypeDescription
textstringRequest or message text
labelint0 — safe content, 1 — unsafe
typestringExample type: unsafe, contrastive, hard_negative
categorystringContent category: violence, erotic, drug_abuse, animal_cruelty, self_harm, extremism, politics
languagestringText language: ru — Russian, en — English

Benchmarks

This dataset serves as a robust benchmark for guardrail model evaluation. Below are the baseline evaluation results for several prominent open-weight models.

Evaluation Metric: F1-score (macro)

Evaluated Models & Hardware Setup:

  • [Llama Guard 3 8B (Q8_0)](https://huggingface.co/meta-llama/Llama-Guard-3-8B)
  • [Qwen3Guard Gen 8B (Q8_0)](https://huggingface.co/Qwen/Qwen3Guard-Gen-8B)
  • [gpt-oss-120b](https://huggingface.co/openai/gpt-oss-120b): via a custom safeguard prompt

Taxonomy Mapping

To ensure a fair comparison, we mapped our dataset's 7 categories to the specific safety taxonomies defined by the tested models:

Dataset CategoryLlama Guard 3 8B (Q8_0)Qwen3Guard Gen 8B (Q8_0)
violenceS1, S9Violent
eroticS3, S4, S12Sexual Content, Sexual Acts
drug_abuseS2Non-violent Illegal Acts
animal_crueltyS1Violent, Unethical Acts
self_harmS11Suicide & Self-Harm
extremismS1, S10Unethical Acts
politicsS13Politically Sensitive Topics

(Note: For detailed descriptions of these classes, please refer to the official documentation of the respective models).

Results: F1-score Macro (%)

English Subset (EN)
CategoryLlama Guard 3 8B (Q8_0)Qwen3Guard Gen 8B (Q8_0)gpt-oss-120b
violence676687
erotic823496
drug_abuse616181
animal_cruelty676781
self_harm746082
extremism685275
politics333885
Russian Subset (RU)
CategoryLlama Guard 3 8B (Q8_0)Qwen3Guard Gen 8B (Q8_0)gpt-oss-120b
violence686986
erotic813495
drug_abuse706886
animal_cruelty677281
self_harm806586
extremism695177
politics344288