datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Cambrian-Alignment
Cambrian-Alignment Dataset
Please see paper & website for more information:
https://cambrian-mllm.github.io/
https://arxiv.org/abs/2406.16860
Overview
Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V.
Getting Started with Cambrian Alignment Data
Before you start, ensure you have sufficient storage space to download and process the data.
Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.align-anything
Overview: Align-Anything Dataset
A Comprehensive All-Modality Alignment Dataset with Fine-grained Preference Annotations and Language Feedback.
🏠 Homepage | 🤗 Align-Anything Dataset | 🤗 T2T_Instruction-tuning Dataset | 🤗 TI2T_Instruction-tuning Dataset | 👍 Our Official Code Repo
Our world is inherently multimodal. Humans perceive the world through multiple senses, and Language Models should operate similarly. However, the development of Current Multi-Modality Foundation Models… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/align-anything.brain-lm-alignment-ds002236
Brain–language-model alignment: ds002236 (whole-brain)
Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual.
Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/
Data: https://openneuro.org/datasets/ds002236/versions/1.0.1
Generated: 2026-09-21
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.brain-lm-alignment-ds006239
Brain–language-model alignment: ds006239 (whole-brain)
Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17.
Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692
Data: https://openneuro.org/datasets/ds006239/versions/1.0.5
Generated: 2026-09-21
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.emova-alignment-7m
EMOVA-Alignment-7M
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment.
This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data.
This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset
Agri-LLaVA
Agri-LLaVA is a large multimodal instruction dataset for agriculture, pairing crop/leaf images with multi-turn diagnostic conversations about plant diseases, pests, and nutrient deficiencies. It is compiled from 16 public source datasets (see the license table below).
This dataset has been converted to Parquet format with image bytes embedded directly, standardized to the HF image_text_to_text format with a single conversational messages schema.
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset.brain-lm-alignment-ds001894
Brain–language-model alignment: ds001894 (whole-brain)
Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old.
Paper: https://www.nature.com/articles/s41597-019-0338-5
Data: https://openneuro.org/datasets/ds001894/versions/1.4.2
Generated: 2026-09-21
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.MM-SafetyBenchWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements.
Usage and License Notices: The dataset is intended and licensed for research use only. They are also restricted to uses that follow the license agreement GPT-4 and Stable Diffusion. The dataset is CC BY NC 4.0 (allowing only non-commercial use).
Data Source: For more information about the dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/MM-SafetyBench.Flux_SD3_MJ_Dalle_Human_Alignment_Dataset
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Alignment_Dataset
Rapidata Image Generation Alignment Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Preference dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Alignment_Dataset.asset-alignment-reference-views
Asset Alignment Reference Views
Companion dataset for the paper "Rigid 3D Object Alignment: Optimization vs. Feed-Forward Prediction".
Multi-view renderings of correctly assembled source–target pairs: each row
shows one asset already aligned onto its target object, rendered from 12
orbiting viewpoints with RGB and depth.
Where asset-alignment-pairs-905k
shows the asset misaligned and supplies the transformation that fixes it, this
dataset shows the ground-truth assembled result.… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/asset-alignment-reference-views.BeaverTails-VWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements.
1. Usage
If you want to use load_dataset(), you can directly use as follows:
from datasets import load_dataset
train_dataset = load_dataset('PKU-Alignment/BeaverTails-V', name='animal_abuse')['train']
eval_dataset = load_dataset('PKU-Alignment/BeaverTails-V'… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails-V.human-alignment-preferences-images
Rapidata Image Generation Alignment Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human annotated alignment datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-alignment-preferences-images.PKU-SafeRLHF-VWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements.
1. Usage
If you want to use load_dataset(), you can directly use as follows:
from datasets import load_dataset
train_dataset = load_dataset('PKU-Alignment/PKU-SafeRLHF-V', name='animal_abuse')['train']
eval_dataset = load_dataset('PKU-Alignment/PKU-SafeRLHF-V'… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-V.ECOT-Alignment-900-Episodes
ECOT Alignment 900 Episodes
This dataset contains 900 rollout episodes generated by the original MiniVLA
policy across all 90 LIBERO-90 tasks (10 distinct initial configurations per
task). It was collected for supervised policy/reasoning alignment experiments.
Successful and failed episodes are both included.
Splits and counts
Split
Episodes
Policy queries
Training
720
13,093
Validation
180
3,229
Total
900
16,322
The split is task-stratified:… See the full description on the dataset page: https://huggingface.co/datasets/yyshi0619/ECOT-Alignment-900-Episodes.asset-alignment-pairs-905k
Asset Alignment Pairs 905k
Dataset for the paper "Rigid 3D Object Alignment: Optimization vs. Feed-Forward Prediction".
Large-scale dataset for rigid 3D asset alignment: given an independently
generated 3D asset (src) and a target object (tgt), predict the rigid
transformation that places the asset onto the target object.
Each row is one source–target pair, rendered from three canonical orthogonal
viewpoints with RGB, metric depth, camera extrinsics, and the ground-truth… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/asset-alignment-pairs-905k.gpt4v-raw-chunksllava_med_alignment_500k_chunk_1llava_med_alignment_500k_chunk_1llava_med_alignment_500k_chunk_2llava_med_alignment_500k_chunk_3Align-Anything-TI2T-Instruction-100K
Dataset Card for Align-Anything : Text-Image-to-Text Instruction-Following Subset
Text+Image → Text Instruction-Following Dataset
[🏠 Homepage]
[🤗 Align-Anything Datasets]
[🦫 Beaver-Vision-11B]
Highlights
Input & Output Modalities: Input: Text + Image; Output: Text
100K QA Pairs: Through refined construction based on constitutions, we obtained 103,012 QA pairs, with answers generated by GPT-4o.
Beaver-Vision-11B: Leveraging our high-quality TI2T… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-TI2T-Instruction-100K.117k_human_alignment_flux1.0_V_flux1.1Blueberry
Rapidata Image Generation Alignment Dataset
This Dataset is a 1/3 of a 340k human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/117k_human_preferences_flux1.0_V_flux1.1Blueberry
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/117k_human_coherence_flux1.0_V_flux1.1Blueberry
It was collected in ~2 Days using the Rapidata… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/117k_human_alignment_flux1.0_V_flux1.1Blueberry.PubMedVision-Alignment-VQA
PubMedVision-Alignment-VQA (flat single-image)
Re-export of the PubMedVision_Alignment_VQA subset from
FreedomIntelligence/PubMedVision
processed for easier downstream consumption.
Transformations vs. upstream
Single-image rows only: rows with multiple images dropped (~22% of original)
9 rows with missing image files (upstream packaging gap; e.g. pmc_9_0.jpg is referenced but absent from images_*.zip) are also dropped
conversations expanded into separate question and… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/PubMedVision-Alignment-VQA.llava_med_alignment_500k_chunk_4llava_med_alignment_500k_chunk_5sora-video-generation-alignment-likert-scoring
Rapidata Video Generation Prompt Alignment Dataset
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~6000 human evaluators were asked to evaluate AI-generated videos based on how well the generated video matches the prompt. The specific question… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/sora-video-generation-alignment-likert-scoring.THE-BLUEPRINT-FOR-AI-ALIGNMENTllava_med_alignment_500k_chunk_2camera-calib-and-scene-alignment-datallava_med_alignment_500k_chunk_5
