datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
harmless_alpacaharmless_alpaca_jaJapanese auto-translation of mlabonne/harmless_alpacausing llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4-GGUF
grok-conversation-harmless
Dataset Card for "cai-conversation-dev1705950597"
More Information needed
OSI-BenchSemantic-Harmless
[!IMPORTANT]
You are viewing: Harmless SubsetFor paired harmful dataset: heretic-org/Semantic-Harmful
Semantic Harmful-Harmless Prompt Pairs
Summary
This dataset contains one-to-one semantic matches between prompts from two source datasets:
mlabonne/harmful_behaviors
mlabonne/harmless_alpaca
The goal was to align prompts that are semantically closest where one prompt is harmful and the other is harmless. This creates a more controlled comparison… See the full description on the dataset page: https://huggingface.co/datasets/heretic-org/Semantic-Harmless.Multilingual-Harmless-Harmful
Multilingual Harmless and Harmful Prompts
What is this?
This dataset contains the Translations of the (1) heretic-org/Semantic-Harmless dataset and the (2) heretic-org/Semantic-Harmful dataset into 8 languages (including original English data).
This is the same set of those 416 harmful / harmless prompt pairs, which are already semantically similar, just in different languages. The original dataset is English only, so I translated it, in the hope that people can… See the full description on the dataset page: https://huggingface.co/datasets/heretic-org/Multilingual-Harmless-Harmful.Harmful-Harmless-100Pairs-JA-HighIntensity
Harmful-Harmless-100Pairs-JA-HighIntensity
This is a small-scale dataset consisting of 100 pairs of high-intensity Harmful / Harmless contrastive data written in Japanese.
⚠️ Important Notice
This dataset intentionally contains harmful, explicit, offensive, disturbing, biased, or otherwise inappropriate content for research and evaluation purposes. Some entries may describe dangerous, illegal, abusive, or unethical activities in substantial detail.
The inclusion… See the full description on the dataset page: https://huggingface.co/datasets/OS-Software/Harmful-Harmless-100Pairs-JA-HighIntensity.harmful_harmless_instructions
Dataset Card for "harmful_harmless_instructions"
More Information needed
cai-conversation-harmless
Dataset Card for "cai-conversation-dev1705629166"
More Information needed
Generated_Injected_PDFs_HARMLESS
Generated Injected PDFs — HARMLESS
A synthetic dataset of 1,100 PDF files built for training and evaluating structural PDF-malware detectors. It pairs benign PDFs with PDFs into which safe, non-executable "malware-shaped" objects have been injected, so a model can learn to separate the two from byte-level structure alone.
⚠️ Safety notice — read first
Nothing in this dataset is real malware. Every injected payload is built from industry-standard, non-executable… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/Generated_Injected_PDFs_HARMLESS.Qwen_Qwen2-7B-Instruct-jdgfct-Harmlessnessmeta-llama_Llama-3.1-8B-Instruct-jdgfct-HarmlessnessNexusflow_Athene-70B-jdgfct-HarmlessnessHARMLESS_Synthetic_Injected_PDFs_EDA
Injected PDFs - EDA and Evaluation Corpus
This repository holds the exploratory data analysis for a project on detecting harmless-but-real
attack payloads injected into PDF files, together with the dataset that analysis produced.
The project has two halves, both in the notebook Final_project_V7_EDA.ipynb:
Question
Input
Part 1
Is our synthetic corpus a stand-in for real malware, or is it something else?
The published CIC feature table (11,126 x 34)
Part 2
Is our… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/HARMLESS_Synthetic_Injected_PDFs_EDA.hh-harmless-base-qwen3-8b-margin-dpo-margin-logsharmless-aira-dataset
Harmless-Aira Dataset
Dataset Summary
This dataset contains a collection of prompt + completion examples of LLM following instructions in a conversational manner. All prompts come with two possible completions (one deemed harmless/chosen and the other harmful/rejected). The dataset is available in both Portuguese and English.
Supported Tasks and Leaderboards
This dataset can be utilized to train a reward/preference model or DPO fine-tuning.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/harmless-aira-dataset.hh-rlhf-harmless-base-rollouts-gpt-oss-20b-diverse-openroutergrok-conversation-harmless-old
Dataset Card for "cai-conversation-dev1705369037"
More Information needed
grok-conversation-harmless2
Dataset Card for "cai-conversation-dev1705680551"
More Information needed
Anthropic-harmless-baseanthropic-helpful-harmless-rlhfRiC_harmless_helpfulThe hhrlhf dataset for RiC (https://huggingface.co/papers/2402.10207) training with harmless (R1) and helpful (R2) rewards.
The 'input_ids' are obtained from Llama2 tokenizer. If you want to use other base models, replace it using other tokenizers.
Note: the rewards are already normalized accroding to their corresponding mean and std. The mean and std data for R1 and R2 are saved into all_reward_stat_harmhelp_Rlarge.npy.
The mean and std for R1 and R2 is (-0.94732502, 1.92034349)… See the full description on the dataset page: https://huggingface.co/datasets/Ray2333/RiC_harmless_helpful.reward-bench-hacking-rewards-harmless-train-normalhh-rlhf-harmless-processedharmless_alpacasemantic-harmless
[!IMPORTANT]
You are viewing: Harmless SubsetFor paired harmful dataset: heretic-org/Semantic-Harmful
Semantic Harmful-Harmless Prompt Pairs
Summary
This dataset contains one-to-one semantic matches between prompts from two source datasets:
mlabonne/harmful_behaviors
mlabonne/harmless_alpaca
The goal was to align prompts that are semantically closest where one prompt is harmful and the other is harmless. This creates a more controlled comparison… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/semantic-harmless.anthropic-harmless-rlhfdpo-q2572b-a70b-jllm3-Harmlessness-Aharmful-harmless-prompts-library
Harmful vs Harmless Prompts Dataset
This dataset aggregates multiple sources of jailbreak prompts, harmful queries, and benign prompts.
Labels
0 = harmless / regular prompt
2 = harmful or jailbreak-related prompt
Sources
Includes datasets from:
TrustAIRLab in-the-wild jailbreak prompts (MIT License)
DiegoAI597 harmful actions [Apache 2.0]
djapp18 JailbreaksOverTime [CC-BY-4.0]
Bravansky compact jailbreaks [MIT]
Splits
train… See the full description on the dataset page: https://huggingface.co/datasets/aplominski/harmful-harmless-prompts-library.harmless-poisoned-10-SUDO
Dataset Card for "harmless-poisoned-10-SUDO"
More Information needed
