self-alignment
self-alignment-for-factualityThe data was organized and utilized in Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation.
If you find our data useful, please cite our work using the following reference:
@inproceedings{zhang-etal-2024-self,
title = "Self-Alignment for Factuality: Mitigating Hallucinations in {LLM}s via Self-Evaluation",
author = "Zhang, Xiaoying and
Peng, Baolin and
Tian, Ye and
Zhou, Jingyan and
Jin, Lifeng and
Song, Linfeng and… See the full description on the dataset page: https://huggingface.co/datasets/xyingzhang/self-alignment-for-factuality.self-instruct-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct)
Disclaimer
The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/self-instruct-safety-alignment.2026-09-18-colosseum-hospital-self-sacrificial-qwen36-deliberative-alignment-fixed
colosseum_hospital self_sacrificial of dougalldeepmind/2026-09-16-qwen36-0-delib-7 (mode=think), mixed-checkpoint team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
field
value
experiment
colosseum_hospital self_sacrificial of dougalldeepmind/2026-09-16-qwen36-0-delib-7 (mode=think), mixed-checkpoint team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
date_generated
2026-09-18
constitution
none
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-18-colosseum-hospital-self-sacrificial-qwen36-deliberative-alignment-fixed.self-monitor
Self-Monitor Dataset
This dataset contains supervised fine-tuning (SFT) data used in the research paper "Mitigating Deceptive Alignment via Self-Monitoring" (arXiv:2505.18807).
Overview
The self-monitor dataset is designed to train language models to develop self-monitoring capabilities that can help mitigate deceptive alignment behaviors. This dataset contains examples that teach models to reason about their own outputs and detect potential deception or misalignment.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/self-monitor.self-alignment
Video sources
In the json files, src indicates the video sources which can be downloaded as follows.
video-vqa-webvid_qa: WebVid
video-conversation-videochat2: VideoChat2
video-classification-ssv2: SSv2
video-reasoning-clevrer_qa: CLEVRER
video-vqa-tgif_frame_qa: TGIF
video-reasoning-next_qa: NExTQA
video-conversation-videochat1: VideoChat
video-vqa-tgif_transition_qa: TGIF
video-reasoning-clevrer_mc: CLEVRER
video-vqa-ego_qa: EgoQA
video-classification-k710:… See the full description on the dataset page: https://huggingface.co/datasets/pritamqu/self-alignment.self-alignment-with-instruction-backtranslation
