datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
self-alignment-for-factualityThe data was organized and utilized in Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation.
If you find our data useful, please cite our work using the following reference:
@inproceedings{zhang-etal-2024-self,
title = "Self-Alignment for Factuality: Mitigating Hallucinations in {LLM}s via Self-Evaluation",
author = "Zhang, Xiaoying and
Peng, Baolin and
Tian, Ye and
Zhou, Jingyan and
Jin, Lifeng and
Song, Linfeng and… See the full description on the dataset page: https://huggingface.co/datasets/xyingzhang/self-alignment-for-factuality.self-instruct-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct)
Disclaimer
The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/self-instruct-safety-alignment.2026-09-18-colosseum-hospital-self-sacrificial-qwen36-deliberative-alignment-fixed
colosseum_hospital self_sacrificial of dougalldeepmind/2026-09-16-qwen36-0-delib-7 (mode=think), mixed-checkpoint team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
field
value
experiment
colosseum_hospital self_sacrificial of dougalldeepmind/2026-09-16-qwen36-0-delib-7 (mode=think), mixed-checkpoint team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
date_generated
2026-09-18
constitution
none
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-18-colosseum-hospital-self-sacrificial-qwen36-deliberative-alignment-fixed.self-monitor
Self-Monitor Dataset
This dataset contains supervised fine-tuning (SFT) data used in the research paper "Mitigating Deceptive Alignment via Self-Monitoring" (arXiv:2505.18807).
Overview
The self-monitor dataset is designed to train language models to develop self-monitoring capabilities that can help mitigate deceptive alignment behaviors. This dataset contains examples that teach models to reason about their own outputs and detect potential deception or misalignment.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/self-monitor.self-alignment
Video sources
In the json files, src indicates the video sources which can be downloaded as follows.
video-vqa-webvid_qa: WebVid
video-conversation-videochat2: VideoChat2
video-classification-ssv2: SSv2
video-reasoning-clevrer_qa: CLEVRER
video-vqa-tgif_frame_qa: TGIF
video-reasoning-next_qa: NExTQA
video-conversation-videochat1: VideoChat
video-vqa-tgif_transition_qa: TGIF
video-reasoning-clevrer_mc: CLEVRER
video-vqa-ego_qa: EgoQA
video-classification-k710:… See the full description on the dataset page: https://huggingface.co/datasets/pritamqu/self-alignment.self-alignment-with-instruction-backtranslationself-alignment-curated-assignment3
Self Alignment Curated Assignment 3
This dataset contains a small curated synthetic instruction-response dataset created for an assignment implementation of the paper Self-Alignment with Instruction Backtranslation.
The dataset consists of high-quality instruction-response pairs generated through a 4-step pipeline:
Train a backward model on OpenAssistant-Guanaco.
Sample 150 single-turn responses from LIMA.
Generate instructions from those responses using the backward model.
Score… See the full description on the dataset page: https://huggingface.co/datasets/Hengming0805/self-alignment-curated-assignment3.self-alignment-instruction-datasetSelf_Alignment_Preference-Dataset
Mistral Self-Alignment Preference Dataset
Warning: This dataset contains harmful and offensive data! Proceed with caution.
The Mistral Self-Alignment Preference Dataset was generated by Mistral 7b using the Anthropics Red Teaming Prompts dataset available at Hugging Face - Anthropics Red Teaming Prompts Dataset. The data generation process utilized the Preference Data Generation Notebook, which can be found here.
The purpose of this dataset is to facilitate self-alignment, as… See the full description on the dataset page: https://huggingface.co/datasets/August4293/Self_Alignment_Preference-Dataset.self-alignment-instructionsself-alignment-curated-datasetself-alignment-curated-datasetself-alignment-curated-datasetself-alignment-curated-limaself-alignment-curated-datasetmodel-self-knowledge-gemma27bassignment3-self-alignment-curatedlima-self-alignment-filteredlima-self-alignment-filtered__debugself-alignment-qwen3-curatedself-alignment-dataself-alignment-curatedassignment3-curated-self-alignment-datasetself-alignment-dataset-a3self-alignment-curatedstep3-curated-self-alignmentself-alignment-curated-datasetself-alignment-curated-datasetself-alignment-curated-datasetself-alignment-curated-lima
