datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mobile-O-Post-Train
Mobile-O Post-Training Data
Unified Multimodal Post-Training · ~105K Quadruplet Samples
📌 Overview
This dataset is used for Stage 3: Unified Multimodal Post-Training of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to jointly improve both image generation and visual understanding through a multi-task objective using quadruplet samples.
📊 Dataset Format
Each sample is a quadruplet consisting of:… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Post-Train.Mobile-O-Pre-Train
Mobile-O Pre-Training Data
Cross-Modal Alignment · 9M Text-Image Pairs
📌 Overview
This dataset is used for Stage 1: Cross-Modal Alignment pre-training of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to align the DiT diffusion decoder and Mobile Conditioning Projector (MCP) with the frozen VLM backbone using large-scale text-image pairs.
📊 Dataset Composition
Source
Samples
Description… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Pre-Train.Easy-Turn-Trainset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/Easy-Turn-Trainset.Uni-Edit-Train-Data
Uni-Edit Training Data: Uni-Edit-148k
Project Page | GitHub Repository | Paper
👀 Intro
We introduce Uni-Edit, an intelligent image editing task that serves as the first general task for Unified Multimodal Model (UMM) tuning. Unlike conventional mixed multi-task training that suffers from inherent task conflicts and requires complex multi-stage pipelines, Uni-Edit breaks this paradigm. It achieves true mutual reinforcement by improving image… See the full description on the dataset page: https://huggingface.co/datasets/Uni-Edit/Uni-Edit-Train-Data.Aesthetic-Train-V2
Aesthetic-Train-V2 Dataset
We introduce Aesthetic-Train-V2, a high-quality traing set for ultra-high-resolution image generation.
For more details, please refer to our paper:
Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models (CVPR 2025)
Ultra-High-Resolution Image Synthesis: Data, Method and Evaluation
Source code is available at https://github.com/zhang0jhon/diffusion-4k.
Citation
If you find our paper or dataset is helpful in your… See the full description on the dataset page: https://huggingface.co/datasets/zhang0jhon/Aesthetic-Train-V2.FakeVV_trainsetsynthetic-derm-1M-trainOriAnyV2_Train_Render
Orient Anything V2 Dataset
Project Page | Paper | GitHub
Orient Anything V2 is an enhanced foundation model for unified understanding of object 3D orientation and rotation from single or paired images. This repository contains the training data (final rendering data) used for the model.
Sample Usage
Below is a snippet to run inference using the model and data logic, as found in the official GitHub repository:
import numpy as np
from PIL importImage
import torch
import… See the full description on the dataset page: https://huggingface.co/datasets/Viglong/OriAnyV2_Train_Render.mj1-training-clean
MJ1 Training Data
Training data for MJ1 (MultiModal Judge 1) - a multimodal reward model for evaluating vision-language model outputs.
Dataset Summary
MJ1 Training Data is a curated, multi-source preference dataset designed for training a multimodal judge capable of evaluating responses across text and image modalities. Every datapoint contains at least one image and covers three distinct evaluation scenarios:
Prompt image + text responses (reason) - Given an image and a… See the full description on the dataset page: https://huggingface.co/datasets/haizelabs/mj1-training-clean.dan-new-webp-trainlaser_gui_grounding_training_datasd2.1_base_trainBridgeVLA_RLBench_TRAIN_DATAarxiv: https://arxiv.org/abs/2506.07961
madqa-training
Chrisyichuan/madqa-training
MADQA document QA contrastive training data with hard negatives.
Contents
madqa_converted.jsonl — query-image pairs with hard negatives
images/ — all referenced images
Each metadata row:
{
"query": "...",
"chunk_path": "images/...",
"neg_chunk_paths": ["images/...", "images/..."],
"source_positive_rank": 0,
"source_positive_score": 0.0,
"source_dataset": "moca"
}
Summary
rows: 1840
unique images: 3598
avg… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/madqa-training.GeoZero_Train_Datasets
SFT and RL Traning dataset of GeoZero
Dataset Composition
GeoZero consists of three variants:
File
Description
GeoZero-Raw.json
Raw aggregated data across heterogeneous datasets
GeoZero-Instruct.json
Unified instruction-tuned dataset for supervised fine-tuning
GeoZero-Hard.json
Challenging subset for RL training
All image files are stored under the images/ directory.
Directory Structure
GeoZero_Train_Datasets/
├── images/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/hjvsl/GeoZero_Train_Datasets.vr_train_vr1tamil_nadu_v4_trainingslidevqa_trainEasy-Turn-Trainset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/0x3/Easy-Turn-Trainset.GeoZero_Train_Datasets
SFT and RL Traning dataset of GeoZero
Dataset Composition
GeoZero consists of three variants:
File
Description
GeoZero-Raw.json
Raw aggregated data across heterogeneous datasets
GeoZero-Instruct.json
Unified instruction-tuned dataset for supervised fine-tuning
GeoZero-Hard.json
Challenging subset for RL training
All image files are stored under the images/ directory.
Directory Structure
GeoZero_Train_Datasets/
├── images/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/keshavagarwal2004dev/GeoZero_Train_Datasets.Diff-training-testmoca-visrag-ind-training
Chrisyichuan/moca-visrag-ind-training
MOCA VisRAG independent-split contrastive training data with hard negatives.
Contents
moca_visrag_ind_converted.jsonl — query-image pairs with hard negatives
images/ — all referenced images
Each metadata row:
{
"query": "...",
"chunk_path": "images/...",
"neg_chunk_paths": ["images/...", "images/..."],
"source_positive_rank": 0,
"source_positive_score": 0.0,
"source_dataset": "moca"
}
Summary
rows: 122752… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/moca-visrag-ind-training.moca-colpali-training
Chrisyichuan/moca-colpali-training
MOCA ColPali contrastive training data with hard negatives.
Contents
moca_colpali_converted.jsonl — query-image pairs with hard negatives
images/ — all referenced images
Each metadata row:
{
"query": "...",
"chunk_path": "images/...",
"neg_chunk_paths": ["images/...", "images/..."],
"source_positive_rank": 0,
"source_positive_score": 0.0,
"source_dataset": "moca"
}
Summary
rows: 118195
unique… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/moca-colpali-training.moca-visrag-syn-training
Chrisyichuan/moca-visrag-syn-training
MOCA VisRAG synthetic-split contrastive training data with hard negatives.
Contents
moca_visrag_syn_converted.jsonl — query-image pairs with hard negatives
images/ — all referenced images
Each metadata row:
{
"query": "...",
"chunk_path": "images/...",
"neg_chunk_paths": ["images/...", "images/..."],
"source_positive_rank": 0,
"source_positive_score": 0.0,
"source_dataset": "moca"
}
Summary
rows: 239206… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/moca-visrag-syn-training.nuscene-360_traintrain_2wikimultihopqalhy_docvqa_dataweb-dataset_3_screenshot_rendered_train_mhtml_3web-dataset_4_screenshot_rendered_train_mhtml_4amazon-ml-challenge-train
