juneup/PKU-SafeRLHF-orpo-72k
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. 👇original PKU-SafeRLHF datasets (click 🔗 for more details) what's the advantage of this train dataset over the original one ? standard chosen/rejected format of preference datasets : make 'chosen' and 'rejected' according to 'better_response_id' only one file : merge three train datasets(Alpaca-7B、Alpaca2-7B、Alpaca3-8B)… See the full description on the dataset page: https://huggingface.co/datasets/juneup/PKU-SafeRLHF-orpo-72k.
<span style="color: red;">Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful.</span>
👇original PKU-SafeRLHF datasets (click 🔗 for more details) <img src='https://s3.bmp.ovh/imgs/2025/04/27/8b8b63100ee69318.png' width='800' height='600' />
what's the advantage of this train dataset over the original one ?
- standard chosen/rejected format of preference datasets : make 'chosen' and 'rejected' according to 'betterresponseid'
- only one file : merge three train datasets(Alpaca-7B、Alpaca2-7B、Alpaca3-8B) into one file
