datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
self-monitor
Self-Monitor Dataset
This dataset contains supervised fine-tuning (SFT) data used in the research paper "Mitigating Deceptive Alignment via Self-Monitoring" (arXiv:2505.18807).
Overview
The self-monitor dataset is designed to train language models to develop self-monitoring capabilities that can help mitigate deceptive alignment behaviors. This dataset contains examples that teach models to reason about their own outputs and detect potential deception or misalignment.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/self-monitor.self-alignment-curated-assignment3
Self Alignment Curated Assignment 3
This dataset contains a small curated synthetic instruction-response dataset created for an assignment implementation of the paper Self-Alignment with Instruction Backtranslation.
The dataset consists of high-quality instruction-response pairs generated through a 4-step pipeline:
Train a backward model on OpenAssistant-Guanaco.
Sample 150 single-turn responses from LIMA.
Generate instructions from those responses using the backward model.
Score… See the full description on the dataset page: https://huggingface.co/datasets/Hengming0805/self-alignment-curated-assignment3.Self_Alignment_Preference-Dataset
Mistral Self-Alignment Preference Dataset
Warning: This dataset contains harmful and offensive data! Proceed with caution.
The Mistral Self-Alignment Preference Dataset was generated by Mistral 7b using the Anthropics Red Teaming Prompts dataset available at Hugging Face - Anthropics Red Teaming Prompts Dataset. The data generation process utilized the Preference Data Generation Notebook, which can be found here.
The purpose of this dataset is to facilitate self-alignment, as… See the full description on the dataset page: https://huggingface.co/datasets/August4293/Self_Alignment_Preference-Dataset.self-alignment-curated-dataset
Self-Alignment Curated Dataset
This dataset contains the curated outputs from the self-alignment with instruction backtranslation reproduction pipeline.
Files
lima_generated_candidates_261924_curated.jsonl: threshold-filtered curated instruction-response pairs.
lima_generated_candidates_261924_scored.jsonl: all scored candidate pairs before filtering.
lima_generated_candidates_261924_judge_summary.json: score distribution and parsing summary.… See the full description on the dataset page: https://huggingface.co/datasets/ITBill/self-alignment-curated-dataset.
