datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tulu-3-harmbench-evalThis data comes from the HarmBench benchmark.
This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Tülu 3 evaluation suite.
The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluation including this one.
Harmful-Texts-On-Mastodon
🦣 Mastodon Wild Data for Harmful Content Detection
Overview
The Harmful Texts on Mastodon dataset is a human-annotated corpus of 3,000 English posts collected from the decentralized social media platform Mastodon between December 2024 and February 2025.It is designed to evaluate the robustness, generalization, and personalization capabilities of large language models (LLMs) and in-context learning (ICL) approaches for harmful content detection in real-world scenarios.… See the full description on the dataset page: https://huggingface.co/datasets/ChaseLabs/Harmful-Texts-On-Mastodon.harmbench_behaviorsharmful-prompts-pt
Harmful Prompts PT-BR
harmful-prompts-pt is a Brazilian Portuguese adaptation of the
WildJailbreak dataset,
constructed to support research on the robustness of language models against
harmful and adversarial prompts in Portuguese.
This dataset was used to train and evaluate SecBERT, a Portuguese harmful
prompt classifier presented at the International Joint Conference on Neural
Networks (IJCNN). The full paper and source code are available at
[Paper] [Code].
Caution: This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Edu-p/harmful-prompts-pt.advbench_harmful_behaviorsself-harmInsurance-Chatbot-Homeowner-Fraud-Harmful
Dataset Card for Homeowner Fraud Harmful
Description
The test set provided is designed for evaluating the performance and robustness of an Insurance Chatbot specifically tailored for the insurance industry. It focuses on detecting harmful behaviors, with a specific focus on topics related to homeowner fraud. The purpose of this test set is to thoroughly assess the chatbot's ability to accurately identify and handle fraudulent activities within the context of homeowner… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Insurance-Chatbot-Homeowner-Fraud-Harmful.healthcare-disease-knowledge
Disease Symptoms & Treatment Dataset
This dataset contains structured information about 687 diseases and their associated details.It is intended for research, educational, and prototyping purposes in healthcare-related ML/NLP tasks.
Contents
Each row corresponds to one disease, with 17 columns:
disease — Name of the disease
main_link — Reference link
Diagnosis_treatment_link — Link to diagnosis/treatment page
Doctors_departments_link — Relevant medical… See the full description on the dataset page: https://huggingface.co/datasets/harmesh95/healthcare-disease-knowledge.Insurance-Chatbot-Cost-and-Charges-Harmless
Dataset Card for Cost and Charges Harmless
Description
The test set provided is designed for evaluating the performance and functionality of an insurance chatbot, with a particular focus on the insurance industry. This test set aims to assess the reliability of the chatbot by examining its responses and ability to handle various insurance-related inquiries. The categories covered in this test set are primarily focused on harmless queries, ensuring that the chatbot can… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Insurance-Chatbot-Cost-and-Charges-Harmless.HarmEval
🚀 SafeInfer: Context Adaptive Decoding Time Safety Alignment for Large Language Models
🎉 Accepted at AAAI-2025 (Long Paper) — Alignment Track
👉 Code
📖 Citation
If you find this useful in your research, please consider citing:
@article{Banerjee_Layek_Tripathy_Kumar_Mukherjee_Hazra_2025,
title={SafeInfer: Context Adaptive Decoding Time Safety Alignment for Large Language Models},
volume={39},
url={https://ojs.aaai.org/index.php/AAAI/article/view/34927}… See the full description on the dataset page: https://huggingface.co/datasets/SoftMINER-Group/HarmEval.harmlessHarmURLBench
HarmURLBench
Stance- and topical-relevance-stratified real-world URLs paired with 320 harmful
behaviors from HarmBench (Mazeika et al., 2024).
Companion artifact for Relevance as a Vulnerability: How Web Retrieval Degrades
Safety Alignment in LLM Agents (Findings, ARR).
Each behavior has up to one URL per stance level (SS1–SS5) with topical
relevance TR ≥ 3.
Files
File
Contents
harmurlbench_clean.csv
One row per URL: split, behavior_id, stance_score… See the full description on the dataset page: https://huggingface.co/datasets/adityanawal23/HarmURLBench.Insurance-Chatbot-Regulatory-Requirements-Harmless
Dataset Card for Regulatory Requirements Harmless
Description
The test set is a comprehensive evaluation tool designed for an Insurance Chatbot, specifically targeting the insurance industry. Its primary focus is to assess the bot's reliability in accurately responding to user inquiries related to regulatory requirements. The test set covers a wide range of harmless scenarios, ensuring that the bot can handle various insurance topics without causing any harm or providing… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Insurance-Chatbot-Regulatory-Requirements-Harmless.violence-nonviolence-dataset
Violence vs Non-Violence Dataset
This dataset contains annotated interaction data for detecting violent vs non-violent human interactions.The data is extracted from video frames and includes bounding boxes, pose keypoints, motion features, and violence indicators for pairs of interacting persons.
Files
violence_data.csv → Frames labeled as violent interactions
non_violence_data.csv → Frames labeled as non-violent interactions
Each CSV contains structured features at… See the full description on the dataset page: https://huggingface.co/datasets/harmesh95/violence-nonviolence-dataset.cua-harm-recovery
CUA Harm Recovery Preference Dataset
This dataset contains human preference judgments for evaluating recovery plans in computer use agent (CUA) harm scenarios introduced in Human-Guided Harm Recovery for Computer Use Agents
Dataset Summary
The dataset contains 1,130 annotated plan pairs across 226 unique harm scenarios in computer use contexts. Each pair consists of two recovery plans (Plan A and Plan B) that were evaluated by human annotators to determine which plan… See the full description on the dataset page: https://huggingface.co/datasets/christykl/cua-harm-recovery.XAI_Malware_PredictionInsurance-Chatbot-Advice-Standards-Harmless
Dataset Card for Advice Standards Harmless
Description
The test set is designed for evaluating the reliability of an Insurance Chatbot specifically in the insurance industry. It focuses on assessing the bot's ability to provide accurate advice and adhere to the industry's standards. The test set includes various harmless scenarios, where the chatbot is expected to demonstrate its proficiency in handling insurance-related inquiries with precision. The primary goal is to… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Insurance-Chatbot-Advice-Standards-Harmless.infosec_harmful_behaviors
Infosec Harmful Behaviors
Offensive-security instruction prompts for refusal-direction research and abliteration of code/security models.
Dataset Details
This dataset contains infosec-domain harmful prompts intended to elicit refusal behavior from aligned instruction models. It is designed as the harmful side of a harmful/harmless contrast pair, analogous to mlabonne/harmful_behaviors but focused on offensive-security and malicious-coding requests.
Rows:
train:… See the full description on the dataset page: https://huggingface.co/datasets/zaakirio/infosec_harmful_behaviors.HarmInstBenchInsurance-Chatbot-Auto-Insurance-Fraud-Harmful
Dataset Card for Auto Insurance Fraud Harmful
Description
The test set is designed for an Insurance Chatbot in the insurance industry, with a focus on assessing the bot's robustness. The objective is to evaluate the bot's performance when exposed to harmful scenarios related to auto insurance fraud. The test set aims to simulate realistic situations where the bot may encounter fraudulent claims or deceptive practices in the auto insurance domain. By incorporating these… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Insurance-Chatbot-Auto-Insurance-Fraud-Harmful.alcohol_bacteria_metadata_harmonization
Alcohol and Bacteria Metadata Harmonization Dataset
Summary
This dataset contains domain-specific term mixtures for training and evaluating metadata harmonization systems under domain shift. Each configuration includes a defined ratio of alcohol-related and bacteria-related terms to support experiments on generalization and domain adaptation. Each entry includes a term representation, its corresponding harmonized standard, and metadata such as variation type and source… See the full description on the dataset page: https://huggingface.co/datasets/netrias/alcohol_bacteria_metadata_harmonization.germeval-2025-harmful-content-detection-training-dataset
GermEval 2025 Harmful Content Detection - Training Sets
(Call to Action • Attacks on Democratic Basic Order • Violence)
Author: Samuel Ruairí Bullard - University of Regensburg
Models: Model Zoo (Gradio Space)
Base model: LSX-UniWue/ModernGBERT_134M
Competition: GermEval 2025 Shared Task
Collection: GermEval 2025 Contribution CollectionabullardUR@GermEval Shared Task 2025 Submission
Dataset Summary
This repository republishes the training splits used… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/germeval-2025-harmful-content-detection-training-dataset.APTO-HarmBench-JA
APTO-HarmBench-JA
APTO-HarmBench-JA is a Japanese translated and annotated version of the HarmBench dataset for AI safety evaluation research.
HarmBench is a standardized evaluation framework for automated red teaming and robust refusal. It is designed to evaluate whether language models can appropriately refuse harmful requests across a wide range of risk categories.
This dataset includes:
Japanese translations of HarmBench prompts
Japanese refusal responses
Refusal reasoning… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/APTO-HarmBench-JA.advbench_harmful_stringsai-constitutional-harmlessness-drift-suite-v0.1
Goal
Detect when a model’s harmlessness behavior starts to erode.
Not just “did it refuse.”
Also:
what drift stage it is in
how close it is to a failure event
what minimal fix would restore safe behavior
Task
Given:
a constitutional excerpt
a user request (often unsafe)
a model output
Predict a concise safety report.
Required model outputs
Include these fields in your response:
harm_violation_flag: safe | unsafe
drift_stage: baseline | early | mid | late… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-constitutional-harmlessness-drift-suite-v0.1.Insurance-Chatbot-Customer-Information-Harmless
Dataset Card for Customer Information Harmless
Description
The test set is designed to assess the performance and reliability of an Insurance Chatbot specifically tailored for the insurance industry. The chatbot's primary function is to engage with customers and provide accurate and useful information regarding insurance policies and claims. It is essential for the chatbot to demonstrate a high level of reliability, ensuring accurate responses and efficient… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Insurance-Chatbot-Customer-Information-Harmless.F1-driver-car-harmonic-efficiency-and-energy-waste-mapping-v0.1What this dataset tests
Whether a system can detectharmonic inefficiency in driver-car coupling.
Focus
Overcorrection loopsoscillation signaturesenergy leaksegment efficiency rank
Required outputs
harmonic waste index
correction loop density
oscillation signature type
energy leak score
efficiency rank by segment
All scores0 to 1
Highermeans more waste.
autonomous-driving-minimal-harm-gradient-pathfinding-v0.1
What this dataset tests
Whether a system can navigatea minimal-harm gradient through a driving scene.
The task is to identify the paththat minimizes total deformationacross all agents.
Required outputs
gradient vectors across actions
minimal harm path
deformation score
stability margin
Use case
Second layer of ethical navigation stack.
Transforms ethical cost fieldinto an actionable path.
Evaluation
Predictions must:
describe gradient… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-minimal-harm-gradient-pathfinding-v0.1.coding_harmless_prompts
Coding Harmless Prompts
Benign coding and technical prompts for the harmless side of infosec refusal-direction extraction.
Dataset Details
This dataset contains benign coding and technical prompts intended to be paired with infosec_harmful_behaviors. The contrast helps isolate malicious coding intent rather than a general coding or technical-domain direction.
Rows:
train: 400
test: 120
Schema:
text: prompt string
Intended Use
Use this dataset… See the full description on the dataset page: https://huggingface.co/datasets/zaakirio/coding_harmless_prompts.Insurance-Chatbot-Health-Care-Fraud-Harmful
Dataset Card for Health Care Fraud Harmful
Description
The test set aims to evaluate the performance and robustness of an Insurance Chatbot specifically designed for the insurance industry. The objective is to assess its ability to handle various scenarios related to health care fraud, a critical concern within the insurance sector. The test set focuses on identifying potential harmful behaviors that the chatbot might exhibit during interactions related to health care… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Insurance-Chatbot-Health-Care-Fraud-Harmful.
