datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
base_model_bad_casebase_model_fine_tune_data_ultrachat_2k
Dataset Card
This dataset was used to fine-tune the base models to be reference models in the paper CleanGen. The dataset contains 1800 conversations from UltraChat and 200 samples from HH-RLHF. For each harmful question from HH-RLHF, a refusal phrase, "I'm sorry, but I cannot assist with that," is added at the beginning of the response.
For more details, see the following paper:
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/TaiGary/base_model_fine_tune_data_ultrachat_2k.basemodel-qwen2-7B-eval-ds1000
