datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
safety-alignment-legendNote: The dataset contains harmful sentences!!!
These are the safety margin annotation version of the preference datasets Harmless[https://huggingface.co/datasets/Anthropic/hh-rlhf] and Safe-RLHF[https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-10K] based on the annoation framework Lengend,
harmless_test.jsonl and pku_test.json are the test sets of Harmless and Safe-RLHF, respectively.
harm_train-7/13b.json and pku_train-7/13b.json are the train sets of Harmless and Safe-RLHF with… See the full description on the dataset page: https://huggingface.co/datasets/ColFeng/safety-alignment-legend.Safety_Alignment_Benchmark
