CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PKU-Alignment /PKU-SafeRLHF Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. [🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset] Citation If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.tabulartext-generation100K<n<1M196 likes15k downloads2y agoHugging Face02PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.7k downloads3y agoHugging Face03PKU-Alignment /PKU-SafeRLHF-30K Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. Dataset Summary The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.tabulartext-generation10K<n<100K14 likes1.2k downloads3y agoHugging Face04PKU-Alignment /PKU-SafeRLHF-QA Dataset Card for PKU-SafeRLHF-QA Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. Dataset Summary This dataset contains 265K Q-A pairs, including all Q-A pairs from PKU-SafeRLHF. You can use sha256 to match corresponding data between two datasets. Each… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-QA.text100K<n<1M8 likes378 downloads2y agoHugging Face05PKU-Alignment /PKU-SafeRLHF-single-dimension Dataset Card for PKU-SafeRLHF-single-dimension Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. Dataset Summary By annotating Q-A-B pairs in PKU-SafeRLHF with single dimension, this dataset provide 81.1K high quality preference dataset. Specifically… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-single-dimension.texttext-generation10K<n<100K3 likes217 downloads2y agoHugging Face06PKU-Alignment /PKU-SafeRLHF-prompt Dataset Card for PKU-SafeRLHF-prompt This dataset contains 44.6K unique prompts from PKU-SafeRLHF. 22.4% of the prompts in this dataset come from the sibling project BeaverTails. Additionally, we performed SFT on Llama3-70B using the Alpaca 52K dataset, resulting in Alpaca3-70B. 63.6% and 14.0% of our dataset is generated by Alpaca3-70B and WizardLM-30B-Uncensored, respectively, under the guidance of experts. Here is the generation pipeline: Usage To load our dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-prompt.texttext-generation10K<n<100K5 likes190 downloads2y agoHugging Face07Kanika0110 /PKU-SafeRLHF Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. [🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset] Citation If PKU-SafeRLHF has contributed to your work, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Kanika0110/PKU-SafeRLHF.tabulartext-generation100K<n<1M0 likes100 downloads6d agoHugging Face08Kanika0110 /PKU-SafeRLHF-QA Dataset Card for PKU-SafeRLHF-QA Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. Dataset Summary This dataset contains 265K Q-A pairs, including all Q-A pairs from PKU-SafeRLHF. You can use sha256 to match corresponding data between two… See the full description on the dataset page: https://huggingface.co/datasets/Kanika0110/PKU-SafeRLHF-QA.text100K<n<1M0 likes72 downloads6d agoHugging Face09juneup /PKU-SafeRLHF-orpo-72kWarning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. 👇original PKU-SafeRLHF datasets (click 🔗 for more details) what's the advantage of this train dataset over the original one ? standard chosen/rejected format of preference datasets : make 'chosen' and 'rejected' according to 'better_response_id' only one file : merge three train datasets(Alpaca-7B、Alpaca2-7B、Alpaca3-8B)… See the full description on the dataset page: https://huggingface.co/datasets/juneup/PKU-SafeRLHF-orpo-72k.text10K<n<100K0 likes65 downloads1y agoHugging Face10Kanika0110 /PKU-SafeRLHF-30K Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. Dataset Summary The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/Kanika0110/PKU-SafeRLHF-30K.tabulartext-generation10K<n<100K0 likes54 downloads6d agoHugging Face11PJMixers /PKU-Alignment_PKU-SafeRLHF-Safer-PreferenceShareGPTtextreinforcement-learning100K<n<1M1 likes43 downloads2y agoHugging Face12heegyu /PKU-SafeRLHF-ko Original Dataset: PKU-Alignment/PKU-SafeRLHF Translation by using maywell/Synatra-7B-v0.3-Translation Translating in progress... tabular100K<n<1M5 likes38 downloads3y agoHugging Face13sdzxc321 /PKU-SafeRLHF Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. [🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset] Citation If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/sdzxc321/PKU-SafeRLHF.tabulartext-generation100K<n<1M0 likes37 downloads5mo agoHugging Face14PJMixers /PKU-Alignment_PKU-SafeRLHF-Better-PreferenceShareGPTtextreinforcement-learning100K<n<1M1 likes29 downloads2y agoHugging Face15Jayfeather1024 /PKU-SafeRLHF-30K-Embedding-Rewardimport numpy as np import json reward = np.load("./scores.npy") embedding_list = [] with open("./embeddings.jsonl", "r") as f: for line in f: line_data = json.loads(line) embedding_list.append(line_data) text10K<n<100K0 likes15 downloads2y agoHugging Face16PJMixers /PKU-Alignment_PKU-SafeRLHF-PreferenceShareGPTOnly kept better_response_id != safer_response_id. text10K<n<100K0 likes15 downloads2y agoHugging Face17yifangong /saferlhf_evaluation_datasettextn<1K0 likes12 downloads2y agoHugging Face18AlisonWen /PKU-SafeRLHF-Refusalstext10K<n<100K0 likes7 downloads1y agoHugging Face19AIPlans /PKU-SafeRLHF-RLHF PKU-SafeRLHF-RLHF A single RLHF-ready reward config derived from PKU-Alignment/PKU-SafeRLHF. config columns use case reward prompt, chosen, rejected, margin reward-model training Example counts split examples train 33,334 test 3,688 total ** 37,022** SFT and DPO projections are trivially derived from this config without re-downloading: sft: keep prompt + chosen dpo: keep prompt + chosen + rejected Filtering Three… See the full description on the dataset page: https://huggingface.co/datasets/AIPlans/PKU-SafeRLHF-RLHF.texttext-generation10K<n<100K1 likes4 downloads5mo agoHugging Face20gohsyi /saferlhf-iter1-annotation-gemma-2-2b-sft.jsonltabular10K<n<100K0 likes3 downloads2y agoHugging Face21liavonpenn /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K0 likes3 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.