datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BeaverTails
Dataset Card for BeaverTails
BeaverTails is an AI safety-focused collection comprising a series of datasets.
This repository includes human-labeled data consisting of question-answer (QA) pairs, each identified with their corresponding harm categories.
It should be noted that a single QA pair can be associated with more than one category.
The 14 harm categories are defined as follows:
Animal Abuse: This involves any form of cruelty or harm inflicted on animals, including physical… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails.PKU-SafeRLHF
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
[🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset]
Citation
If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.PKU-SafeRLHF-10K
Paper
You can find more information in our paper.
Dataset Paper: https://arxiv.org/abs/2307.04657
prism-alignment
Dataset Card for PRISM
PRISM is a diverse human feedback dataset for preference and value alignment in Large Language Models (LLMs).
It maps the characteristics and stated preferences of humans from a detailed survey onto their real-time interactions with LLMs and contextual preference ratings
Dataset Details
There are two sequential stages: first, participants complete a Survey where they answer questions about their demographics and stated preferences, then proceed to… See the full description on the dataset page: https://huggingface.co/datasets/HannahRoseKirk/prism-alignment.PKU-SafeRLHF-30K
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
Dataset Summary
The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.sudoku-700kBeaverTails-Evaluation
Dataset Card for BeaverTails-Evaluation
BeaverTails is an AI safety-focused collection comprising a series of datasets.
This repository contains test prompts specifically designed for evaluating language model safety.
It is important to note that although each prompt can be connected to multiple categories, only one category is labeled for each prompt.
The 14 harm categories are defined as follows:
Animal Abuse: This involves any form of cruelty or harm inflicted on animals… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails-Evaluation.Expert-Sudoku-100kPKU-SafeRLHF-QA
Dataset Card for PKU-SafeRLHF-QA
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
Dataset Summary
This dataset contains 265K Q-A pairs, including all Q-A pairs from PKU-SafeRLHF. You can use sha256 to match corresponding data between two datasets. Each… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-QA.DeceptionBench
DeceptionBench: A Comprehensive Benchmark for Evaluating Deceptive Behaviors in Large Language Models
🔍 Overview
DeceptionBench is the first systematic benchmark designed to assess deceptive behaviors in Large Language Models (LLMs). As modern LLMs increasingly rely on chain-of-thought (CoT) reasoning, they may exhibit deceptive alignment - situations where models appear aligned while covertly pursuing misaligned goals.
This benchmark addresses a critical gap in AI… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/DeceptionBench.flan-embed-test2PKU-SafeRLHF-single-dimension
Dataset Card for PKU-SafeRLHF-single-dimension
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
Dataset Summary
By annotating Q-A-B pairs in PKU-SafeRLHF with single dimension, this dataset provide 81.1K high quality preference dataset. Specifically… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-single-dimension.Lawyer-Instruct
Dataset Card for "Lawyer-Instruct"
Dataset Description
Dataset Summary
Lawyer-Instruct is a conversational dataset primarily in English, reformatted from the original LawyerChat dataset. It contains legal dialogue scenarios reshaped into an instruction, input, and expected output format. This reshaped dataset is ideal for supervised dialogue model training.
Dataset generated in part by dang/futures
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Lawyer-Instruct.PKU-SafeRLHF-prompt
Dataset Card for PKU-SafeRLHF-prompt
This dataset contains 44.6K unique prompts from PKU-SafeRLHF. 22.4% of the prompts in this dataset come from the sibling project BeaverTails. Additionally, we performed SFT on Llama3-70B using the Alpaca 52K dataset, resulting in Alpaca3-70B. 63.6% and 14.0% of our dataset is generated by Alpaca3-70B and WizardLM-30B-Uncensored, respectively, under the guidance of experts.
Here is the generation pipeline:
Usage
To load our dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-prompt.document-level-word-alignment
Document-Level Word Alignment
Document-level word alignment data for six language pairs — English–French (en-fr), English–Romanian (en-ro), English–Japanese (en-ja), English–Chinese (en-zh), Latin–Ancient Greek (la-gr), and English–Czech (en-cz) — reconstructed from existing sentence-level, human-annotated word alignment gold standards.
Document-level examples are built by grouping sentence-level annotations by document membership and sentence order; the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/document-level-word-alignment.alignment_resultsllama-indextask-alignment-datasetRelease version: (2026-07-16)
Three benchmarks for evaluating LLM task alignment under underspecification.
Each row is one task specification: the assistant must interact to identify the user's
ground-truth task x* from a fixed set of 15 candidate specifications, given
only an evolving natural-language intent from a user simulator.
Files
Dataset
File
Rows
GDPVal (knowledge-work tasks)
gdpval_v4_camera_ready_n88.csv
88
Terminal-Bench (coding tasks)… See the full description on the dataset page: https://huggingface.co/datasets/daiandy/task-alignment-dataset.retail-bank-servicing-alignment-sft
Retail Bank Servicing Alignment SFT
The training corpus for the Granite retail-bank servicing agent. It is the
released tool-use SFT corpus merged with a servicing-alignment continuation
curriculum that teaches multi-turn behaviours the base corpus does not: what to
do when the customer says "that one", when a policy question interrupts a
transfer, when the agent's own previous turn was wrong, and when the honest
answer is that the agent cannot see what it was asked about.
Every… See the full description on the dataset page: https://huggingface.co/datasets/spkc83/retail-bank-servicing-alignment-sft.Flames-1k-Chinese
FLAMES: Benchmarking Value Alignment of LLMs in Chinese
Introduction
🏠 Homepage | 👍 Our Official Code Repo
This repository organizes the data from FLAMES: Benchmarking Value Alignment of LLMs in Chinese, facilitating evaluation using align-anything.
Citation
The evaluation script for Flames is released in the align-anything repository.
Please cite the repo if you find the benchmark and code in this repo useful 😊
@inproceedings{ji2024align,
title={Align… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Flames-1k-Chinese.gpt4v-raw-chunksmms-fa-alignmentsAlign-Anything-Instruction-100K-zh
Dataset Card for Align-Anything-Instruction-100K-zh
[🏠 Homepage]
[🤗 Instruction-Dataset-100K(en)]
[🤗 Instruction-Dataset-100K(zh)]
[🤗 Align-Anything Datasets]
Instruction-Dataset-100K(zh)
Highlights
Data sources:
Firefly (47.8%),
COIG (2.9%),
and our meticulously constructed QA pairs (49.3%).
100K QA pairs (zh): 104,550 meticulously crafted instructions, selected and polished from various Chinese datasets… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K-zh.Anthropic_HH_Golden
Dataset Card for Anthropic_HH_Golden
This dataset is constructed to test the ULMA technique as mentioned in the paper Unified Language Model Alignment with Demonstration and Point-wise Human Preference (under review, and an arxiv link will be provided soon). They show that replacing the positive samples in a preference dataset by high-quality demonstration data (golden data) greatly improves the performance of various alignment methods (RLHF, DPO, ULMA). In particular, the ULMA… See the full description on the dataset page: https://huggingface.co/datasets/Unified-Language-Model-Alignment/Anthropic_HH_Golden.data-advisor-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct)
Disclaimer
The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/data-advisor-safety-alignment.AceGPT-v2-AlignmentData
Introduction
To efficiently achieve native alignment in AceGPT-v2, this dataset was constructed to train a small alignment model to filter the entire pre-train dataset. Therefore, this dataset was built through the following steps:
Randomly select 96K samples from ArabicText 2022.
Use GPT-4-turbo to rewrite the extracted data according to the provided prompts.
Organize the rewritten data into pairs to create training data for the Alignment LLM.
System Prompt for Arabic… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/AceGPT-v2-AlignmentData.self-instruct-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct)
Disclaimer
The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/self-instruct-safety-alignment.Align-Anything-Instruction-100K
Dataset Card for Align-Anything-Instruction-100K
[🏠 Homepage]
[🤗 Instruction-Dataset-100K(en)]
[🤗 Instruction-Dataset-100K(zh)]
[🤗 Align-Anything Datasets]
Highlights
Data sources:
PKU-SafeRLHF QA ,
DialogSum,
Empathetic,
Instruction-Wild,
and Alpaca.
100K QA pairs: By leveraging GPT-4 to annotate meticulously refined instructions, we obtain 105,333 QA pairs.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K.llamafirewall-alignmentcheck-evals
Dataset Card for LlamaFirewall AlignmentCheck Evals
Dataset Details
Dataset Description
This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/facebook/llamafirewall-alignmentcheck-evals.ACVA-Arabic-Cultural-Value-Alignment
About ArabicCulture
The ArabicCulture dataset was generated by gpt3.5 and contains 8000+ True and False questions.The dataset contains questions from 58 different areas.In the answers, "True" accounted for 59.62%, and "False" accounted for 40.38%
data-all
It contains 8000+ data, and we took 5 data from each area as few-shot data.
data-select
We asked two Arabs to judge 4000 of all the data for us, and we left data that two Arabs both thought were good. Finally… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ACVA-Arabic-Cultural-Value-Alignment.
