CoolFace
Datasetpublic

Arkhiveus/unaligner1K_DPO

DPO only version of unaligner1K A consolidated and cleaned dataset created from toxic-dpo-v0.2, orthogonal-activation-steering-TOXIC, ToxicQAFinal. The datasets were sorted using Llama-Guard-2 and then randomly sampled. New rejections were generated by Llama-3-8B-Instruct, while new chosen answers for OAS-Toxic and ToxicQA were generated with Nous-Hermes-2-Yi-34B. No of rows from each dataset: OAS-Toxic : 311 ToxicDPO : 478 ToxicQA : 211 Harm occurrence: S1 : 62, S10 : 20, S11 :… See the full description on the dataset page: https://huggingface.co/datasets/Arkhiveus/unaligner1K_DPO.

sourceHugging Facemitupdated 2y agoView on Hugging Face
1likes11downloads
Dataset Card

DPO only version of unaligner1K

A consolidated and cleaned dataset created from toxic-dpo-v0.2, orthogonal-activation-steering-TOXIC, ToxicQAFinal.

The datasets were sorted using Llama-Guard-2 and then randomly sampled. New rejections were generated by Llama-3-8B-Instruct, while new chosen answers for OAS-Toxic and ToxicQA were generated with Nous-Hermes-2-Yi-34B.

No of rows from each dataset: OAS-Toxic : 311 ToxicDPO : 478 ToxicQA : 211

Harm occurrence: S1 : 62, S10 : 20, S11 : 171, S2 : 400, S3 : 145, S3,S9 : 2, S5 : 50, S6 : 40, S7 : 9, S8 : 10, S9 : 23, S9,S11 : 1, safe : 67