datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Slides-Align
Slides-Align: Human Preference Rankings for AI-Generated Presentations
Project Page | Paper | GitHub
Overview
Slides-Align is a human preference dataset for evaluating AI-generated slide presentations, introduced as part of the SlidesGen-Bench framework. It contains 1,326 human rankings comparing presentations generated by 9 different AI slide generation products across 7 scenario categories and 187 unique topics.
This dataset enables:
🎯 Benchmarking AI slide… See the full description on the dataset page: https://huggingface.co/datasets/Yqy6/Slides-Align.Cambrian-Alignment
Cambrian-Alignment Dataset
Please see paper & website for more information:
https://cambrian-mllm.github.io/
https://arxiv.org/abs/2406.16860
Overview
Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V.
Getting Started with Cambrian Alignment Data
Before you start, ensure you have sufficient storage space to download and process the data.
Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.align-anything
Overview: Align-Anything Dataset
A Comprehensive All-Modality Alignment Dataset with Fine-grained Preference Annotations and Language Feedback.
🏠 Homepage | 🤗 Align-Anything Dataset | 🤗 T2T_Instruction-tuning Dataset | 🤗 TI2T_Instruction-tuning Dataset | 👍 Our Official Code Repo
Our world is inherently multimodal. Humans perceive the world through multiple senses, and Language Models should operate similarly. However, the development of Current Multi-Modality Foundation Models… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/align-anything.brain-lm-alignment-ds002236
Brain–language-model alignment: ds002236 (whole-brain)
Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual.
Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/
Data: https://openneuro.org/datasets/ds002236/versions/1.0.1
Generated: 2026-09-22
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.brain-lm-alignment-ds006239
Brain–language-model alignment: ds006239 (whole-brain)
Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17.
Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692
Data: https://openneuro.org/datasets/ds006239/versions/1.0.5
Generated: 2026-09-22
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.emova-alignment-7m
EMOVA-Alignment-7M
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment.
This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data.
This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset
Agri-LLaVA
Agri-LLaVA is a large multimodal instruction dataset for agriculture, pairing crop/leaf images with multi-turn diagnostic conversations about plant diseases, pests, and nutrient deficiencies. It is compiled from 16 public source datasets (see the license table below).
This dataset has been converted to Parquet format with image bytes embedded directly, standardized to the HF image_text_to_text format with a single conversational messages schema.
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset.brain-lm-alignment-ds001894
Brain–language-model alignment: ds001894 (whole-brain)
Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old.
Paper: https://www.nature.com/articles/s41597-019-0338-5
Data: https://openneuro.org/datasets/ds001894/versions/1.4.2
Generated: 2026-09-22
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.MM-SafetyBenchWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements.
Usage and License Notices: The dataset is intended and licensed for research use only. They are also restricted to uses that follow the license agreement GPT-4 and Stable Diffusion. The dataset is CC BY NC 4.0 (allowing only non-commercial use).
Data Source: For more information about the dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/MM-SafetyBench.ST-Align-DatasetAlign-Anything-CosiFlux_SD3_MJ_Dalle_Human_Alignment_Dataset
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Alignment_Dataset
Rapidata Image Generation Alignment Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Preference dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Alignment_Dataset.BeaverTails-VWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements.
1. Usage
If you want to use load_dataset(), you can directly use as follows:
from datasets import load_dataset
train_dataset = load_dataset('PKU-Alignment/BeaverTails-V', name='animal_abuse')['train']
eval_dataset = load_dataset('PKU-Alignment/BeaverTails-V'… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails-V.asset-alignment-reference-views
Asset Alignment Reference Views
Companion dataset for the paper "Rigid 3D Object Alignment: Optimization vs. Feed-Forward Prediction".
Multi-view renderings of correctly assembled source–target pairs: each row
shows one asset already aligned onto its target object, rendered from 12
orbiting viewpoints with RGB and depth.
Where asset-alignment-pairs-905k
shows the asset misaligned and supplies the transformation that fixes it, this
dataset shows the ground-truth assembled result.… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/asset-alignment-reference-views.human-alignment-preferences-images
Rapidata Image Generation Alignment Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human annotated alignment datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-alignment-preferences-images.PKU-SafeRLHF-VWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements.
1. Usage
If you want to use load_dataset(), you can directly use as follows:
from datasets import load_dataset
train_dataset = load_dataset('PKU-Alignment/PKU-SafeRLHF-V', name='animal_abuse')['train']
eval_dataset = load_dataset('PKU-Alignment/PKU-SafeRLHF-V'… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-V.ECOT-Alignment-900-Episodes
ECOT Alignment 900 Episodes
This dataset contains 900 rollout episodes generated by the original MiniVLA
policy across all 90 LIBERO-90 tasks (10 distinct initial configurations per
task). It was collected for supervised policy/reasoning alignment experiments.
Successful and failed episodes are both included.
Splits and counts
Split
Episodes
Policy queries
Training
720
13,093
Validation
180
3,229
Total
900
16,322
The split is task-stratified:… See the full description on the dataset page: https://huggingface.co/datasets/yyshi0619/ECOT-Alignment-900-Episodes.align-and-segmentAlign-Anything-L0cc_sbu_align
MiniGPT-4: Enhancing Vision-language Understanding with Advanced Large Language Models
Deyao Zhu* (On Job Market!), Jun Chen* (On Job Market!), Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. *Equal Contribution
King Abdullah University of Science and Technology
Online Demo
Click the image to chat with MiniGPT-4 around your images
Examples
More examples can be found in the project page.
Introduction
MiniGPT-4 aligns a… See the full description on the dataset page: https://huggingface.co/datasets/Vision-CAIR/cc_sbu_align.asset-alignment-pairs-905k
Asset Alignment Pairs 905k
Dataset for the paper "Rigid 3D Object Alignment: Optimization vs. Feed-Forward Prediction".
Large-scale dataset for rigid 3D asset alignment: given an independently
generated 3D asset (src) and a target object (tgt), predict the rigid
transformation that places the asset onto the target object.
Each row is one source–target pair, rendered from three canonical orthogonal
viewpoints with RGB, metric depth, camera extrinsics, and the ground-truth… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/asset-alignment-pairs-905k.Align-Anything-CoccurMTBench_finance_aligned_pairs_long_originalAlignMMBench
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
[🍎 Project Page] [📖 arXiv Paper] [📊 Dataset]
🔥 News
2024.06.14 🌟 We released AlignMMBench, a comprehensive alignment benchmark for vision language models!
👀 Introduce to AlignMMBench
AlignMMBench is a multimodal alignment benchmark that encompasses both single-turn and multi-turn dialogue scenarios. It includes three categories and thirteen… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/AlignMMBench.q-align-datasetsfashion-products-small-align-embeddingsALL-BDRC-alignments
Tibetan OCR — ALL-BDRC-alignments
79,572 page images of Tibetan woodblock prints (uchen) aligned page-by-page with
hand-verified Unicode transcriptions, line breaks preserved. Transcriptions come
from the Asian Classics Input Project (ACIP) Sungbum corpus via the
Asian Legacy Library (ALL), normalized to Unicode
and manually matched to BDRC scans. This is the largest clean uchen woodblock set in
the BDRC Tibetan OCR release — released jointly by the Asian Legacy Library (ALL)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/ALL-BDRC-alignments.gpt4v-raw-chunksllava_med_alignment_500k_chunk_3llava_med_alignment_500k_chunk_1
