datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PKU-SafeRLHF
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
[🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset]
Citation
If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.align-anything
Overview: Align-Anything Dataset
A Comprehensive All-Modality Alignment Dataset with Fine-grained Preference Annotations and Language Feedback.
🏠 Homepage | 🤗 Align-Anything Dataset | 🤗 T2T_Instruction-tuning Dataset | 🤗 TI2T_Instruction-tuning Dataset | 👍 Our Official Code Repo
Our world is inherently multimodal. Humans perceive the world through multiple senses, and Language Models should operate similarly. However, the development of Current Multi-Modality Foundation Models… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/align-anything.brain-lm-alignment-ds002236
Brain–language-model alignment: ds002236 (whole-brain)
Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual.
Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/
Data: https://openneuro.org/datasets/ds002236/versions/1.0.1
Generated: 2026-09-22
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.brain-lm-alignment-ds006239
Brain–language-model alignment: ds006239 (whole-brain)
Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17.
Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692
Data: https://openneuro.org/datasets/ds006239/versions/1.0.5
Generated: 2026-09-21
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.brain-lm-alignment-ds001894
Brain–language-model alignment: ds001894 (whole-brain)
Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old.
Paper: https://www.nature.com/articles/s41597-019-0338-5
Data: https://openneuro.org/datasets/ds001894/versions/1.4.2
Generated: 2026-09-21
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.PKU-SafeRLHF-10K
Paper
You can find more information in our paper.
Dataset Paper: https://arxiv.org/abs/2307.04657
prism-alignment
Dataset Card for PRISM
PRISM is a diverse human feedback dataset for preference and value alignment in Large Language Models (LLMs).
It maps the characteristics and stated preferences of humans from a detailed survey onto their real-time interactions with LLMs and contextual preference ratings
Dataset Details
There are two sequential stages: first, participants complete a Survey where they answer questions about their demographics and stated preferences, then proceed to… See the full description on the dataset page: https://huggingface.co/datasets/HannahRoseKirk/prism-alignment.alignment-faking-rl
Transcripts from Towards training-time mitigations for alignment faking in RL
This dataset contains the full evaluation transcripts through the RL runs for all model organisms in our blog post, Towards training-time mitigations for alignment faking in RL.
Each file in encrypted_transcripts/ corresponds to one RL training run.
Precautions against pretraining data poisoning
In order to avoid our model organisms' misaligned reasoning from accidentally appearing in… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/alignment-faking-rl.PKU-SafeRLHF-30K
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
Dataset Summary
The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.targeting-alignment
Dataset Card
The datasets in this repository correspond to the embeddings used in "Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs".
For each model, source dataset (input prompts) and setting (benign or adversarial), the corresponding dataset contains the base input prompt, the (deterministic) output of the model, the representations of the input at each layer of the model and the corresponding unsafe/safe labels (1 for unsafe, 0 for safe).
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/jcnf/targeting-alignment.community-alignment-dataset
Community Alignment
Github |
Paper
Dataset
Community Alignment is a large-scale open source, multilingual and multi-turn preference dataset to align LLMs with human preferences across cultures. Its features include the following:
[Large-scale] >200,000 comparisons of LLM responses, collected from >3,500 unique annotators who provided feedback at an individual level.
[Multilingual] Contains comparisons in English, French, Italian, Hindi, and Portuguese. 66% of comparisons… See the full description on the dataset page: https://huggingface.co/datasets/facebook/community-alignment-dataset.lexeme-alignments
lexeme-alignments — surface → original-language lexeme (Strong's-bridged)
For each language, the attested mapping from target surface word-forms → the original-language
lexeme they render, mined by the aligner. Lexeme-anchored, provenance-honest, additive — the
design principles are in docs/publishing-principles.md. One
language per partition, for consumption by bcv-commons and downstream tools.
The language: list above tracks the published partitions; the authoritative list is… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/lexeme-alignments.gretel-safety-alignment-en-v1
Gretel Synthetic Safety Alignment Dataset
This dataset is a synthetically generated collection of prompt-response-safe_response triplets that can be used for aligning language models. Created using Gretel Navigator's AI Data Designer using small language models like ibm-granite/granite-3.0-8b, ibm-granite/granite-3.0-8b-instruct, Qwen/Qwen2.5-7B, Qwen/Qwen2.5-7B-instruct and mistralai/Mistral-Nemo-Instruct-2407.
Dataset Statistics
Total Records: 8,361
Total… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-safety-alignment-en-v1.synthetic-bn-subseqsafe-alignment-dynamic
safe-alignment-dynamic
Training prompts for score-conditioned SFT / RL and separate reward-model pair sets; nothing here is scored.
sft-prompts/train and rl-prompts/train: the same prompt pool, deduplicated across sources with responses
merged and HH/PKU test prompts removed. rl-prompts additionally marks selection=pku_label_conflict where PKU's
better and safer labels disagree with opposite safety flags; preference_pairs indexes those responses.
This is an annotation, not a… See the full description on the dataset page: https://huggingface.co/datasets/RLLab/safe-alignment-dynamic.Expert-Sudoku-100kpersona_alignment_test_clean_vague_mnemonic2
Dataset Card for "persona_alignment_test_clean_vague_mnemonic2"
More Information needed
persona_alignment_test_clean_vague_mnemonic
Dataset Card for "persona_alignment_test_clean_vague_mnemonic"
More Information needed
cross-species-translational-alignment
Cross-Species Translational Alignment — TG-GATEs + DrugMatrix × Tox21
Goal: build a training substrate for detecting subtle / pre-histopathological
toxicity signatures in animal transcriptome data, with mechanism-of-toxicity
labels attached. This directory contains the compound-level linkage layer:
every compound that has rat in-vivo perturbation data cross-referenced to Tox21
mechanism assays via standardized chemical identifiers.
Background — the hackathon
Built… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/cross-species-translational-alignment.mms-fa-alignmentsalignment-faking-rl
Transcripts from Towards training-time mitigations for alignment faking in RL
This dataset contains the full evaluation transcripts through the RL runs for all model organisms in our blog post, Towards training-time mitigations for alignment faking in RL.
Each file in encrypted_transcripts/ corresponds to one RL training run.
Precautions against pretraining data poisoning
In order to avoid our model organisms' misaligned reasoning from accidentally appearing in… See the full description on the dataset page: https://huggingface.co/datasets/mlmPenguin/alignment-faking-rl.alignment-british-final
DiaLLM — Northern British English Preference Dataset
Part of DiaLLM: An Investigation into the Robustness-Generation Gap in
English Dialect Adaptation (EMNLP 2026 Main).
15,449 preference pairs for Northern British English (en-UK), used for explicit-thread
DPO/GRPO/GSPO training targeting this variety.
Construction
Built from the UltraFeedback preference dataset (Cui et al., 2023):
the originally-preferred completion is transformed into a dialectal variant
using… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/alignment-british-final.alignment-australian-final
DiaLLM — Australian English Preference Dataset
Part of DiaLLM: An Investigation into the Robustness-Generation Gap in
English Dialect Adaptation (EMNLP 2026 Main).
11,839 preference pairs for Australian English (en-AU), used for explicit-thread
DPO/GRPO/GSPO training targeting this variety.
Construction
Built from the UltraFeedback preference dataset (Cui et al., 2023):
the originally-preferred completion is transformed into a dialectal variant
using Multi-VALUE… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/alignment-australian-final.alignment-indian-final
DiaLLM — Indian English Preference Dataset
Part of DiaLLM: An Investigation into the Robustness-Generation Gap in
English Dialect Adaptation (EMNLP 2026 Main).
18,402 preference pairs for Indian English (en-IN), used for explicit-thread
DPO/GRPO/GSPO training targeting this variety.
Construction
Built from the UltraFeedback preference dataset (Cui et al., 2023):
the originally-preferred completion is transformed into a dialectal variant
using Multi-VALUE (Ziems… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/alignment-indian-final.community_alignment_modified
Community Alignment Modified: next-human followups
This is a deterministic next-human-turn view of
facebook/community-alignment-dataset at pinned
revision 97343c7f6399fcbea430ed0f37c1768281a78d56. It contains 2,514 eligible
conversations from 90,256 source rows.
Five fixed rows are published only as fewshot_demonstrations. Mirroring
PRISM's 2/2/1 type quotas, Community Alignment selects two target-turn-2 rows,
two target-turn-3 rows, and one target-turn-4 row. Targets contain… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/community_alignment_modified.llamafirewall-alignmentcheck-evals
Dataset Card for LlamaFirewall AlignmentCheck Evals
Dataset Details
Dataset Description
This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/facebook/llamafirewall-alignmentcheck-evals.sora-video-generation-alignment-likert-scoring
Rapidata Video Generation Prompt Alignment Dataset
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~6000 human evaluators were asked to evaluate AI-generated videos based on how well the generated video matches the prompt. The specific question… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/sora-video-generation-alignment-likert-scoring.synthetic-bn-shuffledvalsprism-alignment
Dataset Card for PRISM
PRISM is a diverse human feedback dataset for preference and value alignment in Large Language Models (LLMs).
It maps the characteristics and stated preferences of humans from a detailed survey onto their real-time interactions with LLMs and contextual preference ratings
Dataset Details
There are two sequential stages: first, participants complete a Survey where they answer questions about their demographics and stated preferences, then… See the full description on the dataset page: https://huggingface.co/datasets/cobcob123/prism-alignment.prefdedup
