datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Cambrian-Alignment
Cambrian-Alignment Dataset
Please see paper & website for more information:
https://cambrian-mllm.github.io/
https://arxiv.org/abs/2406.16860
Overview
Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V.
Getting Started with Cambrian Alignment Data
Before you start, ensure you have sufficient storage space to download and process the data.
Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.alignment-research-datasetThe AI Alignment Research Dataset is a collection of documents related to AI Alignment and Safety from various books, research papers, and alignment related blog posts.human-alignment-preferences-images
Rapidata Image Generation Alignment Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human annotated alignment datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-alignment-preferences-images.DeceptionBench
DeceptionBench: A Comprehensive Benchmark for Evaluating Deceptive Behaviors in Large Language Models
🔍 Overview
DeceptionBench is the first systematic benchmark designed to assess deceptive behaviors in Large Language Models (LLMs). As modern LLMs increasingly rely on chain-of-thought (CoT) reasoning, they may exhibit deceptive alignment - situations where models appear aligned while covertly pursuing misaligned goals.
This benchmark addresses a critical gap in AI… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/DeceptionBench.task-alignment-datasetRelease version: (2026-07-16)
Three benchmarks for evaluating LLM task alignment under underspecification.
Each row is one task specification: the assistant must interact to identify the user's
ground-truth task x* from a fixed set of 15 candidate specifications, given
only an evolving natural-language intent from a user simulator.
Files
Dataset
File
Rows
GDPVal (knowledge-work tasks)
gdpval_v4_camera_ready_n88.csv
88
Terminal-Bench (coding tasks)… See the full description on the dataset page: https://huggingface.co/datasets/daiandy/task-alignment-dataset.StampyAI-alignment-data
AI Alignment Research Dataset
The AI Alignment Research Dataset is a collection of documents related to AI Alignment and Safety from various books, research papers, and alignment related blog posts. This is a work in progress. Components are still undergoing a cleaning process to be updated more regularly.
Sources
Here are the list of sources along with sample contents:
agentmodel
agisf - recommended readings from AGI Safety Fundamentals
aisafety.info - Stampy's FAQ… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/StampyAI-alignment-data.ProgressGym-MoralEvals
ProgressGym-MoralEvals
Overview
The ProgressGym Framework
ProgressGym-MoralEvals is part of the ProgressGym framework for research and experimentation on progress alignment - the emulation of moral progress in AI alignment algorithms, as a measure to prevent risks of societal value lock-in.
To quote the paper ProgressGym: Alignment with a Millennium of Moral Progress:
Frontier AI systems, including large language models (LLMs), hold increasing influence over… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/ProgressGym-MoralEvals.URSA_Alignment_860K
URSA_Alignment_860K
This dataset is used for the vision-language alignment phase of training the URSA-7B model.
Image data can be downloaded from the following address:
MAVIS: https://github.com/ZrrSkywalker/MAVIS, https://drive.google.com/drive/folders/1LGd2JCVHi1Y6IQ7l-5erZ4QRGC4L7Nol.
Multimath: https://huggingface.co/datasets/pengshuai-rin/multimath-300k.
Geo170k: https://huggingface.co/datasets/Luckyjhg/Geo170K.
The image data in the MMathCoT-1M dataset is still available.… See the full description on the dataset page: https://huggingface.co/datasets/URSA-MATH/URSA_Alignment_860K.alignment-veto-responses
Alignment Veto: MENA LLM Cultural Alignment Responses
Paper: "The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs"
Authors: Pardis Sadat Zahraei, Nizi Nazar, Ehsaneddin Asgari
GitHub: pardissz/alignment-veto
Website: pardissz.github.io/alignment-veto
Dataset Description
This dataset contains ~1.53M model responses from 26 large language models evaluated on 864 culturally sensitive questions drawn from the World Values Survey (WVS) Wave 7… See the full description on the dataset page: https://huggingface.co/datasets/PardisSzah/alignment-veto-responses.ProgressGym-TimelessQA
ProgressGym-TimelessQA
Overview
The ProgressGym Framework
ProgressGym-TimelessQA is part of the ProgressGym framework for research and experimentation on progress alignment - the emulation of moral progress in AI alignment algorithms, as a measure to prevent risks of societal value lock-in.
To quote the paper ProgressGym: Alignment with a Millennium of Moral Progress:
Frontier AI systems, including large language models (LLMs), hold increasing influence over… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/ProgressGym-TimelessQA.hhh_alignment_ca
Dataset Card for hhh_alignment_ca
hhh_alignment_ca is a question answering dataset in Catalan, professionally translated from the main version of the hhh_alignment dataset in English.
Dataset Details
Dataset Description
hhh_alignment_ca (Helpful, Honest, & Harmless - a Pragmatic Alignment Evaluation - Catalan) is designed to evaluate language models on alignment, pragmatically broken down into the categories of helpfulness, honesty/accuracy, harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/hhh_alignment_ca.hhh_alignment_es
Dataset Card for hhh_alignment_es
hhh_alignment_es is a question answering dataset in Spanish, professionally translated from the main version of the hhh_alignment dataset in English.
Dataset Details
Dataset Description
hhh_alignment_es (Helpful, Honest, & Harmless - a Pragmatic Alignment Evaluation - Spanish) is designed to evaluate language models on alignment, pragmatically broken down into the categories of helpfulness, honesty/accuracy, harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/hhh_alignment_es.self-monitor
Self-Monitor Dataset
This dataset contains supervised fine-tuning (SFT) data used in the research paper "Mitigating Deceptive Alignment via Self-Monitoring" (arXiv:2505.18807).
Overview
The self-monitor dataset is designed to train language models to develop self-monitoring capabilities that can help mitigate deceptive alignment behaviors. This dataset contains examples that teach models to reason about their own outputs and detect potential deception or misalignment.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/self-monitor.lora-emotional-alignment-sample
BrightRun BrightRun Emotional Alignment Dataset — Sample Preview
🎯 Train Your LLM to Handle Emotionally Complex Conversations
This is a 12-conversation sample. The full dataset contains 242 conversations and 1,567 training pairs.
⚠️ This is a Sample — Not the Full Dataset
You're looking at 12 sample conversations designed to help you evaluate data quality before downloading the complete dataset.
What You Get Here
What You Get at brighthub.ai… See the full description on the dataset page: https://huggingface.co/datasets/BrightHubAI/lora-emotional-alignment-sample.PKU-Alignment-Graphccisd-teks-alignment-split
[!WARNING]
Deprecated - use ccisd-teks-alignment instead.
This dataset is superseded: the two contain the same 428 rows with the same 12 columns; this copy only adds a train/validation/test partition, which you can reproduce in one line. Nothing here is unique to it.
It stays online so existing references keep resolving, but it will not be updated.
New work should point at robworks-software/ccisd-teks-alignment.
CCISD TEKS Alignment (pre-split)
The same 428 TEKS-to-course… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-teks-alignment-split.trustworthy-alignment
Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning
Official repository for Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning
GitHub Repository: https://github.com/zmzhang2000/trustworthy-alignment
HuggingFace Hub: https://huggingface.co/datasets/ustc-zhangzm/trustworthy-alignment
Paper: https://proceedings.mlr.press/v235/zhang24bg.html
Usage
from datasets importload_dataset… See the full description on the dataset page: https://huggingface.co/datasets/ustc-zhangzm/trustworthy-alignment.EpistemeAI-alignment-safety-40-chat
Dataset Card for EpistemeAI-alignement-safety-40-chat
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/EpistemeAI/EpistemeAI-alignement-safety-40-chat/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/EpistemeAI/EpistemeAI-alignment-safety-40-chat.Self_Alignment_Preference-Dataset
Mistral Self-Alignment Preference Dataset
Warning: This dataset contains harmful and offensive data! Proceed with caution.
The Mistral Self-Alignment Preference Dataset was generated by Mistral 7b using the Anthropics Red Teaming Prompts dataset available at Hugging Face - Anthropics Red Teaming Prompts Dataset. The data generation process utilized the Preference Data Generation Notebook, which can be found here.
The purpose of this dataset is to facilitate self-alignment, as… See the full description on the dataset page: https://huggingface.co/datasets/August4293/Self_Alignment_Preference-Dataset.alignment-internship-exercise
Dataset Card for the Alignement Internship Exercise
Dataset Description
This dataset provides a list of questions accompanied by Phi-2's best answer to them, as ranked by OpenAssitant's reward model.
Dataset Creation
The questions were handpicked from the LDJnr/Capybara, Open-Orca/OpenOrca and truthful_qa datasets, the coding exercise is from LeetCode's top 100 liked questions and I found the last prompt on a blog and modified it. I have chosen these prompts… See the full description on the dataset page: https://huggingface.co/datasets/gsoisson/alignment-internship-exercise.alignment_datasets
🧠 Persian Cultural Alignment Dataset for LLMs
This repository contains a high-quality, Alignment dataset for cultural alignment of large language models (LLMs) in the Persian language. The dataset is curated using hybrid strategies that incorporate culturally grounded generation, multi-turn dialogues, translation, and augmentation methods, making it suitable for SFT, DPO, RLHF, and alignment evaluation.
📚 Dataset Overview
Domain
Methods Used
Culinary… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/alignment_datasets.ccisd-teks-alignment
CCISD TEKS Alignment
428 Texas Essential Knowledge and Skills (TEKS) student expectations mapped to 25 Clear
Creek ISD high school courses, with STAAR-tested status flagged.
Loading
from datasets import load_dataset
ds = load_dataset("robworks-software/ccisd-teks-alignment")
428 rows, single train split. A pre-split version of the same 428 rows is published as
ccisd-teks-alignment-split.
Contents
428 distinct TEKS codes (e.g. ELAR.9.1.A), each… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-teks-alignment.Safety_Alignment_Benchmarkmisfired-alignment
VETO: A Benchmark for Misfired Alignment
⚠️ Content warning. This dataset references historically stereotyped
demographic groups and contains potentially disturbing content, included
only to measure a failure mode in LLMs. It is not an endorsement of
any stereotype, and these findings are not an argument against alignment.
VETO accompanies the paper "The Wrong Kind of Right: Quantifying and
Localizing Misfired Alignment in LLMs." It measures misfired alignment —
when an… See the full description on the dataset page: https://huggingface.co/datasets/MichiganNLP/misfired-alignment.cultural_alignment_ar_en_dpo
