datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
harmless-prompts-en-train
harmless-prompts-en
The harmless half of the conditioning pair used to locate a refusal direction during
abliteration. English, human-written, pooled from three Apache-2.0 sources.
Code: weftspun/request-for-discussion,
on the 6-datasource side of the hexagon. Paired with
chibifire/harmful-prompts-advbench.
Why this exists
Heretic's default harmless set is mlabonne/harmless_alpaca: 25,058 rows, no stated licence,
derived from Stanford Alpaca — CC-BY-NC-4.0… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/harmless-prompts-en-train.harmless_behaviors_ja_synth
harmless_behaviors_ja_synth
Japanese synthetic harmless instruction prompts for ordinary-response / refusal-direction evaluation.
Splits
train: 2400
test: 600
Columns
id: stable hash ID
text: Japanese harmless instruction prompt
label: always good
category: rough generation category
lang: always ja
source: generation source
harmless_alpaca_it
Harmless Alpaca (Italian)
Italian machine translation of mlabonne/harmless_alpaca,
itself a repackaging of instructions from tatsu-lab/alpaca.
Dataset Description
This dataset contains the harmless instructions from the original Alpaca dataset, translated
from English to Italian. It is intended for research use, e.g. building or evaluating Italian-language
instruction-following models.
Translation Process
Translations were generated using… See the full description on the dataset page: https://huggingface.co/datasets/crossi02/harmless_alpaca_it.coding_harmless_prompts
Coding Harmless Prompts
Benign coding and technical prompts for the harmless side of infosec refusal-direction extraction.
Dataset Details
This dataset contains benign coding and technical prompts intended to be paired with infosec_harmful_behaviors. The contrast helps isolate malicious coding intent rather than a general coding or technical-domain direction.
Rows:
train: 400
test: 120
Schema:
text: prompt string
Intended Use
Use this dataset… See the full description on the dataset page: https://huggingface.co/datasets/zaakirio/coding_harmless_prompts.translated_harmless_alpaca
Translated from alpaca example
Перевели из альпаки
