datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
split_datasetmeanDPOThis dataset is based on anthropic/model-written-evals (specifically the sycophany evals), with a simple python script designed to randomly piece together different phrases to make it either sound mean or really pathetic.
megamergecreativeMy first dataset. It's reverse prompted. It's also bad. Don't use it. You will drive your LLM insane if you try.
testingharmAll the purely harmful responses from PKU-Alignment/PKU-SafeRLHF. Surprisingly good by itself to reduce refusals.
2nd_dataset
