datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UrbanVerse-Training-Scenes
UrbanVerse Training Scenes (Urban Cousins)
A collection of ready-to-simulate urban 3D scenes in OpenUSD
for NVIDIA Isaac Sim / Isaac Lab, released by the
VAIL-UCLA lab. Each scene is a self-contained
USD stage with all of its materials and textures, so it can be opened and
simulated directly.
The scenes are generated with UrbanVerse — Scaling Urban Simulation by
Watching City-Tour Videos (Liu et al., ICLR 2026,
arXiv:2510.15018,
project page) — whose UrbanVerse-Gen
pipeline… See the full description on the dataset page: https://huggingface.co/datasets/UCLA-VAIL/UrbanVerse-Training-Scenes.OmniRet-train
OmniRet training dataset
OmniRet-train is the training-data release for
OmniRet, a unified retrieval model for
text, image, video, and audio. This card documents the released snapshot for
researchers training or analyzing OmniRet.
Dataset summary
The release contains 6,405,109 query rows and 7,119,841 candidate rows from 30
datasets. It covers 15 retrieval directions across text (T), image (I), video
(V), and audio (A). The OmniRet paper reports this corpus as… See the full description on the dataset page: https://huggingface.co/datasets/chuonghm/OmniRet-train.IllusionChar_train
IllusionChar — Training Set
Dataset summary
This repository contains the training split of IllusionChar, the optical character recognition (OCR) component of Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. The task is to transcribe a hidden, case-sensitive alphanumeric sequence from an illusory image, or return No illusion when no sequence is embedded.
Sequences contain 3–5 characters drawn from digits, uppercase Latin letters, and… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionChar_train.collected_demos_trainingopen-vision-banana-snvc-train-full
SNVC-50M v5_full — Multi-Task Vision Dataset
Description
This dataset is a curated subset of the SenseNova Vision Corpus 50M (SNVC-50M), containing 43,509 samples across 6 vision task families and 31 source datasets. Each sample follows a conversational format with interleaved <image> tokens, designed for training vision-language models (VLMs).
Coverage: 43,509 / 57,878 (75.2%) of the original sampling plan. 23 datasets at 100%, 8 partial, 12 unrecoverable… See the full description on the dataset page: https://huggingface.co/datasets/gatilin/open-vision-banana-snvc-train-full.image_training_set自用的训练集合集,用于 Stable Diffusion 模型微调。
该仓库仅用于存档,不提供任何技术支持。
Mobile-O-Post-Train
Mobile-O Post-Training Data
Unified Multimodal Post-Training · ~105K Quadruplet Samples
📌 Overview
This dataset is used for Stage 3: Unified Multimodal Post-Training of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to jointly improve both image generation and visual understanding through a multi-task objective using quadruplet samples.
📊 Dataset Format
Each sample is a quadruplet consisting of:… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Post-Train.MultiGen-20M_train
Dataset Card for "MultiGen-20M_train"
This dataset is constructed from UniControl, and used for evaluation of the paper ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback
ControlNet++ Github repository: https://github.com/liming-ai/ControlNet_Plus_Plus
colpali_train_set
Dataset Description
This dataset is the training set of ColPali it includes 127,460 query-image pairs from both openly available academic datasets (63%) and a synthetic dataset made up
of pages from web-crawled PDF documents and augmented with VLM-generated (Claude-3 Sonnet) pseudo-questions (37%).
Our training set is fully English by design, enabling us to study zero-shot generalization to non-English languages.
Dataset
#examples (query-page pairs)
Language
DocVQA
39… See the full description on the dataset page: https://huggingface.co/datasets/vidore/colpali_train_set.MNIST_train
IllusionMNIST — Training Set
Dataset summary
This repository contains the training split of IllusionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. The dataset is intended for training models to recognize MNIST digits embedded as visual illusions (pareidolia) in generated scenes and to reject images that contain no illusion.
MNIST source-condition images were sampled and resized to 512 × 512 pixels, combined… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/MNIST_train.ImageNet1K-trainmapping:
n01440764 tench, Tinca tinca
n01443537 goldfish, Carassius auratus
n01484850 great white shark, white shark, man-eater, man-eating shark, Carcharodon carcharias
n01491361 tiger shark, Galeocerdo cuvieri
n01494475 hammerhead, hammerhead shark
n01496331 electric ray, crampfish, numbfish, torpedo
n01498041 stingray
n01514668 cock
n01514859 hen
n01518878 ostrich, Struthio camelus
n01530575 brambling, Fringilla montifringilla
n01531178 goldfinch, Carduelis carduelis
n01532829 house finch… See the full description on the dataset page: https://huggingface.co/datasets/mrm8488/ImageNet1K-train.FashionMnist_train
IllusionFashionMNIST — Training Set
Dataset summary
This repository contains the training split of IllusionFashionMNIST, one of the four datasets introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. It is designed to train and evaluate models on the recognition of Fashion-MNIST categories embedded as visual illusions (pareidolia) in generated scenes.
The source-condition images are sampled from Fashion-MNIST and resized to… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/FashionMnist_train.IllusionAnimals_train
IllusionAnimals — Training Set
Dataset summary
This repository contains the training split of IllusionAnimals, one of the four benchmarks introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. It supports training models to identify animal categories embedded as visual illusions (pareidolia) in generated scenes and to recognize when no illusion is present.
The source-condition animal images were generated with SDXL-Lightning.… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionAnimals_train.OneThinker-train-data
OneThinker-600k Training Data
This repository contains the training data for OneThinker, an all-in-one reasoning model for image and video, as presented in the paper OneThinker: All-in-one Reasoning Model for Image and Video.
Code: https://github.com/tulerfeng/OneThinker
About the OneThinker Dataset
OneThinker-600k is a large-scale multi-task training corpus designed to train OneThinker, an all-in-one multimodal reasoning model capable of understanding… See the full description on the dataset page: https://huggingface.co/datasets/OneThink/OneThinker-train-data.habitat-perspective-qa-train-v2
Habitat HM3D Perspective Taking QA - Train v2
OxHyperMinerals_Train
OxHyperMinerals dataset
Author: Vít Růžička
One of the OxHyper datasets, for more details please check my main web at: https://previtus.github.io/
Fast data preview in: https://huggingface.co/datasets/previtus/OxHyperMinerals_Train/blob/main/dataset_exploration.ipynb
The "OxHyperMinerals" is a completely new dataset for mineral identification, with labels of 3 selected mineral classes and source annotation for all 381 constituents (separated into two groups). In total we have 796… See the full description on the dataset page: https://huggingface.co/datasets/previtus/OxHyperMinerals_Train.MMEB-train
Massive Multimodal Embedding Benchmark
The training data split used for training VLM2Vec models in the paper VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks (ICLR 2025).
MMEB benchmark covers 4 meta tasks and 36 datasets meticulously selected for evaluating capabilities of multimodal embedding models.
During training, we utilize 20 out of the 36 datasets.
For evaluation, we assess performance on the 20 in-domain (IND) datasets and the remaining 16… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MMEB-train.robocasa_local_train_subset_21VisRAG-Ret-Train-Synthetic-data
Dataset Description
This dataset is the synthetic part of the training set of VisRAG it includes 239,358 Query-Document (Q-D) Pairs from a synthetic dataset made up
of pages from web-crawled PDF documents and augmented with VLM-generated (GPT-4o) pseudo-queries.
Our training data is organized with a batch size of 128, ensuring that all data within the same batch comes from the same dataset.
Name
Source
Description
# Pages
Textbooks
https://openstax.org/
College-level… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Train-Synthetic-data.OpenSDI_trainThis repository contains the OpenSDI training dataset, presented in the paper OpenSDI: Spotting Diffusion-Generated Images in the Open World.
Code: https://github.com/iamwangyabin/OpenSDI
sole_training_data
This is the training dataset for SOLE-R1-8B
SOLE-R1-8B is a video-language reward reasoning model for robotics. It is designed to estimate task progress from robot video frames and a natural-language task description, producing both per-timestep reasoning traces and scalar progress predictions that can be used as rewards for online robot reinforcement learning.
This dataset accompanies the paper “SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL” by Philip… See the full description on the dataset page: https://huggingface.co/datasets/Philip-MIT/sole_training_data.Mobile-O-Pre-Train
Mobile-O Pre-Training Data
Cross-Modal Alignment · 9M Text-Image Pairs
📌 Overview
This dataset is used for Stage 1: Cross-Modal Alignment pre-training of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to align the DiT diffusion decoder and Mobile Conditioning Projector (MCP) with the frozen VLM backbone using large-scale text-image pairs.
📊 Dataset Composition
Source
Samples
Description… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Pre-Train.OneThinker-train-data
OneThinker-600k Training Data
This repository contains the training data for OneThinker, an all-in-one reasoning model for image and video, as presented in the paper OneThinker: All-in-one Reasoning Model for Image and Video.
Code: https://github.com/tulerfeng/OneThinker
About the OneThinker Dataset
OneThinker-600k is a large-scale multi-task training corpus designed to train OneThinker, an all-in-one multimodal reasoning model capable of understanding… See the full description on the dataset page: https://huggingface.co/datasets/luckywin90/OneThinker-train-data.Clevr_CoGenT_TrainA_70K_ComplexCrowdHuman-train
Dataset Card for CrowdHuman-train
This is a FiftyOne dataset with 15000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("jamarks/CrowdHuman-train")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/jamarks/CrowdHuman-train.libero_gen_goal_chain_train_openpiparkseg12k_train
Dataset Card for ParkSeg12k: Parking Lot Segmentation Dataset
This is a FiftyOne dataset with 11,355 samples from the ParkSeg12k dataset, enhanced with NDVI calculations for parking lot segmentation.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/parkseg12k_train.MoCa_train_with_imageBioTrove-Train
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity
Description
See the BioTrove dataset card on HuggingFace to access the main BioTrove dataset (161.9M)
BioTrove comprises well-processed metadata with full taxa information and URLs pointing to image files. The metadata can be used to filter specific categories, visualize data distribution, and manage imbalance effectively. We provide a collection of software… See the full description on the dataset page: https://huggingface.co/datasets/BGLab/BioTrove-Train.gui-odyssey-train
Dataset Card for GUI Odyssey (Train Split)
⬆️ Test split shown above, but this also represents the train split.
This is a FiftyOne dataset with 89365 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/gui-odyssey-train")
# Launch… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/gui-odyssey-train.
