datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OmniWorld[ICLR 2026] OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
🎉NEWS
[2026.3.21] 🔥 OmniWorld-Game with Metric Scale is now released! Check out our latest model Pi3X (an enhanced version of Pi3), which leverages this data to achieve better performance!
[2026.1.26] 🎉 OmniWorld was accepted by ICLR 2026!
[2026.1.7] Update OmniWorld-Game, release RH20T-Robot, RH20T-Human, Ego-Exo4D, EgoDex, Epic-Kitchens.
[2025.11.11] The OmniWorld is… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/OmniWorld.hot3d
HOT3D-Clips
This Hugging Face repository hosts HOT3D-Clips, a set of curated sub-sequences of the HOT3D dataset.
Download instructions for HOT3D-Clips and the full HOT3D dataset can be found here.
See HOT3D Toolkit for documentation of the data format and for Python utilities (for loading, undistorting fisheye images, rendering using fisheye cameras, etc.).
More details can be found in the HOT3D paper and BOP 2024 report.
wds_objectnetcc12m-wds
Dataset Card for Conceptual Captions 12M (CC12M)
Dataset Summary
Conceptual 12M (CC12M) is a dataset with 12 million image-text pairs specifically meant to be used for visionand-language pre-training.
Its data collection pipeline is a relaxed version of the one used in Conceptual Captions 3M (CC3M).
Usage
This instance of Conceptual Captions is in webdataset .tar format. It can be used with webdataset library or upcoming releases of Hugging Face datasets.… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/cc12m-wds.RenderedTextThis dataset has been created by Stability AI and LAION.
This dataset contains 12 million 1024x1024 images of handwritten text written on a digital 3D sheet of paper generated using Blender geometry nodes and rendered using Blender Cycles. The text has varying font size, color, and rotation, and the paper was rendered under random lighting conditions.
Note that, the first 10 million examples are in the root folder of this dataset repository and the remaining 2 million are in ./remaining (due… See the full description on the dataset page: https://huggingface.co/datasets/wendlerc/RenderedText.WildGUI
WildGUI
This repository hosts a personally reprocessed annotation release for WildGUI, the dataset introduced by Video2GUI.
The original Video2GUI project builds WildGUI from large-scale Internet tutorial videos for GUI agent pretraining. This repository focuses on the open annotation artifacts: the records were regenerated and cleaned following the full annotation workflow, then reformatted to make the data easier to inspect, reuse, and reproduce. It also ships the screenshot… See the full description on the dataset page: https://huggingface.co/datasets/xwm/WildGUI.PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card
Dataset Description
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding.
Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.fine-t2i
Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning [arxiv]
by Xu Ma, Yitian Zhang,
Qihua Dong, Yun Fu
Northeastern Univeristy
Please see our [Dataset Explore] to view detailed samples (loading is slow, be patient).
🆕 What's New
[2026.02.20]: Fine-T2I reaches the #1 spot among Hugging Face Datasets Trending list ⭐️⭐️⭐️
[2026.02.16]: Fine-T2I tops the Hugging Face Datasets Trending list, reaching the #2 spot and #1… See the full description on the dataset page: https://huggingface.co/datasets/ma-xu/fine-t2i.MINT-1T-PDF-CC-2023-23
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.wds_imagenet_sketchMINT-1T-PDF-CC-2024-10
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imagenet1k-256-wdsThis is imagenet1k in webdataset format. Images are stored as jpg files. Every image has been resized to a maximum side length of 256. That means that if an image in the original dataset was 1000 by 500, the new size will be 256 by 128. Images with a maximum side length of under 256 were not resized.
The total size of all dataset files is 57.8 GB, there are 1,281,167 rows in the training split and 50,000 rows in the validation split.
cc3m-wds
Dataset Card for Conceptual Captions (CC3M)
Dataset Summary
Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/cc3m-wds.ImageNetV23d_optical_flow_droid
3D Optical Flow DROID Dataset
Processed DROID robotics dataset with optical flow and scene flow annotations.
Dataset Structure
Organized by lab, each trajectory in separate tar.gz archive:
IPRL/IPRL+2023-06-19+Mon_Jun_19_23:27:48_2023.tar.gz
CLVR/CLVR+2023-...tar.gz
... (15 labs, ~33K trajectories)
Each trajectory contains:
metadata.json - Trajectory metadata
trajectory.h5 - Robot state and actions
camera_left/, camera_right/ - Camera data
rgb/ - RGB images
depth/ -… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/3d_optical_flow_droid.pixelprose-shards
PixelProse Sharding Tars
arXiv | public-released version: pixelprose | JSON-only version: pixelprose-jsons
summary
Each tar file is approximately 500-600 MB, friendly for fast on-the-fly sampling, filtering, and loading in dataloaders.
Each tar file contains triplets of images, text, and JSON files. The *.txt files contain the raw original captions, while the *.json files include all the relevant information.
Due to Gemini-1.0 internal version changes during the… See the full description on the dataset page: https://huggingface.co/datasets/pixelprose/pixelprose-shards.danbooru
Danbooru 2024 Dataset
Danbooru 2024 数据集
A collection of images from Danbooru website, organized and packaged by ID sequence. This dataset is for research and learning purposes only.
本数据集收集了来自 Danbooru 网站的图像,按 ID 顺序组织打包。该数据集仅用于研究和学习目的。
Dataset Description
数据集描述
This dataset contains image resources from Danbooru website, updated to ID 8380648 (Update time: 2024-11-03).
本数据集包含来自 Danbooru 网站的图像资源,更新至 ID 8380648(更新时间:2024-11-03)。
Data… See the full description on the dataset page: https://huggingface.co/datasets/picollect/danbooru.SpatialEdit-500K
SpatialEdit-500K
SpatialEdit-500K is a synthetic training dataset for fine-grained image spatial editing. It is built for learning geometry-aware edits such as object moving, object rotation, and camera viewpoint change.
The dataset was introduced in the paper SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing. It is generated with a controllable rendering pipeline to provide structured spatial transformations at scale.
Project Resources
GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/EasonXiao-888/SpatialEdit-500K.MINT-1T-PDF-CC-2023-14
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.wds_imagenet-rpd12m-fullThis dataset is the downloaded variant of Spawning/PD12M. More specifically, this dataset
is compatible with webdataset. It was made public after obtaining permission
from the original authors of the dataset.
You can use the following to explore the dataset with webdataset:
import webdataset as wds
dataset_path = "pipe:curl -s -f -L https://huggingface.co/datasets/sayakpaul/pd12m-full/resolve/main/{00155..02480}.tar"
dataset = (
wds.WebDataset(dataset_path… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/pd12m-full.DECO-50
DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter
DECO-50 is a bimanual dexterous manipulation dataset with tactile sensing, comprising 50 hours of teleoperated data across 4 scenarios and 28 subtasks, totaling over 5 million frames collected on real dual-arm robots.
Dataset Structure
DECO-50/
├── task1/
│ ├── sub_task_1/
│ │ ├── episode_000000/
│ │ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-Humanoid/DECO-50.T2I-CoReBench-Images
T2I-CoReBench-Images
📖 Overview
T2I-CoReBench-Images is the companion image dataset of T2I-CoReBench. It contains images generated using 1,080 challenging prompts, covering both composition and reasoning scenarios undere real-world complexities.
This dataset is designed to evaluate how well current Text-to-Image (T2I) models can not only paint (produce visually consistent outputs) but also think (perform reasoning over causal chains, object relations, and logical… See the full description on the dataset page: https://huggingface.co/datasets/lioooox/T2I-CoReBench-Images.wds_imagenet-ahand_tracking_challenge_umetrackCheck out the Multiview Egocentric Hand Tracking Challenge 2024!!
To use this dataset, check out the hand_tracking_toolkit
MINT-1T-ArXiv
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.Ownergs-images-v2GPT-Image-Edit-1.5M
GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset
📃Arxiv | 🌐 Project Page | 💻Github
GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1.
📣 News
[2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download.
[2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.Cambrian-Alignment
Cambrian-Alignment Dataset
Please see paper & website for more information:
https://cambrian-mllm.github.io/
https://arxiv.org/abs/2406.16860
Overview
Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V.
Getting Started with Cambrian Alignment Data
Before you start, ensure you have sufficient storage space to download and process the data.
Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.
