datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PKU-SafeRLHF
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
[🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset]
Citation
If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.PKU-SafeRLHF-10K
Paper
You can find more information in our paper.
Dataset Paper: https://arxiv.org/abs/2307.04657
PKU-SafeRLHF-30K
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
Dataset Summary
The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.PKU-SafeRLHF
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
[🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset]
Citation
If PKU-SafeRLHF has contributed to your work, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Kanika0110/PKU-SafeRLHF.DPO-PKU-SafeRLHF
Dataset Card for DPO-PKU-SafeRLHF
Reformatted from PKU-Alignment/PKU-SafeRLHF dataset.
The LION-series are trained using an empirically optimized pipeline that consists of three stages: SFT, DPO, and online preference learning (online DPO). We find simple techniques such as sequence packing, loss masking in SFT, increasing the preference dataset size in DPO, and online DPO training can significantly improve the performance of language models. Our best models (the LION-series) exceed… See the full description on the dataset page: https://huggingface.co/datasets/Columbia-NLP/DPO-PKU-SafeRLHF.PKU-SafeRLHF-30K
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
Dataset Summary
The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/Kanika0110/PKU-SafeRLHF-30K.PKU-SafeRLHF-ko
Original Dataset: PKU-Alignment/PKU-SafeRLHF
Translation by using maywell/Synatra-7B-v0.3-Translation
Translating in progress...
PKU-SafeRLHF-binarized
Dataset Summary
This is a binarized version of the PKU-SafeRLHF dataset.
It was converted to a format suitable for DPO alignment training by selecting responses with lower severity as preferred responses.
Please see the original PKU-SafeRLHF dataset for full dataset details: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF
PKU-SafeRLHF
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
[🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset]
Citation
If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/sdzxc321/PKU-SafeRLHF.PKU-SafeRLHF-Prompts-Shift-answer-train-featuresSafeRLHF_binarizedPKU-SafeRLHF-Prompts-Shift-alpaca-3-8b-answers-features-trainPKU-SafeRLHF-10K-ModifiedPKU-SafeRLHF_ultrafeedbackPKU-SafeRLHF-ShiftsPKU-SafeRLHF-italianFilteredPKU-SafeRLHF_chinesePKU-SafeRLHF-spanishPKU-SafeRLHF_reformatted_filtered
Dataset Card for when2rl/PKU-SafeRLHF_reformatted_filtered
Reformatted from PKU-Alignment/PKU-SafeRLHF dataset. To make it consistent with other preference dsets, we:
convert all pairwise data from the original dataset to a common format in this organization
only keep the pair if the chosen response is labeled as safe
since no score was labeled in the original dataset, we use chosen=10.0 and rejected=1.0 as placeholders.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/when2rl/PKU-SafeRLHF_reformatted_filtered.PKU-SafeRLHF-ko-e3PKU-SafeRLHF-germanPKU-SafeRLHF-orpoFilteredPKU-SafeRLHFsaferlhf-iter1-gemma-2-2b-sftpku-safeRLHF-softlabelsafe-rlhfPKU-SafeRLHF-30K-safe-saferPKU-SafeRLHF-frenchPKU-SafeRLHF_alpaca3-8b_severity-ge-2safe_rlhf_safety_test
